Skip to content
heliaAOT
HELIA HUB

Measure a conversion

A conversion is worth measuring on the Ambiq silicon it will run on. heliaPROFILER (hpx) builds profiler firmware around a model, flashes a target, captures what the run cost, and writes a result directory you can compare against another one.

First validate the module with a trusted oracle and appropriate tolerance. You need a profiler-supported board, toolchain and probe configured for that board. For power, complete the profiler’s monitor and wiring setup as well. These are separate prerequisites from successfully converting the model.

Keep the model file and its hash, target, clock, engine versions, compiler settings and placement with the result. The commands below operate hardware; a successful conversion alone does not produce a measurement bundle.

hpx profile measures whole-inference timing, optional per-layer measurements and the firmware image it builds. Available counters depend on the processor. Power and energy require a supported monitor, correct wiring and the corresponding measurement mode. Every option is in the hpx profile reference. The engine is one of the things it varies, so the same model can be measured through heliaRT, heliaAOT or TFLM. Keep the model, board, clock, toolchain and measurement settings aligned, then inspect the differences recorded in the bundles.

Terminal window
pip install 'helia-profiler[aot]'
hpx doctor

hpx doctor checks installed host tools and Python dependencies. It does not verify a connected probe or board, and its default checks do not establish AOT readiness. Install the aot extra and follow the board setup instructions. See the heliaPROFILER documentation for the full setup. heliaPROFILER has its own release and configuration contract; record the installed version and follow that version’s reference if it differs from the recipe below. An AOT package version does not pin the profiler, runtime engine or toolchain for you.

The recipe below is heliaPROFILER’s engine comparison example, adapted here with matching model, board and profiling settings. Replace my_model.tflite with your model path. The engine configuration and output directory differ; generated code and runtime build modes can differ too.

hpx_rt.yml
model:
path: my_model.tflite
arena_size: 131072
engine:
type: helia-rt
config:
variant: release-with-logs
target:
board: apollo510_evb
toolchain: arm-none-eabi-gcc
profiling:
pmu_counters:
cpu: default
per_layer: true
iterations: 5
warmup: 2
output:
format: csv
dir: ./results/comparison_rt
detailed: true
hpx_aot.yml
model:
path: my_model.tflite
arena_size: 131072
engine:
type: helia-aot
config:
prefix: hpx
module_name: hpx_model
target:
board: apollo510_evb
toolchain: arm-none-eabi-gcc
profiling:
pmu_counters:
cpu: default
per_layer: true
iterations: 5
warmup: 2
output:
format: csv
dir: ./results/comparison_aot
detailed: true

Each successful run writes the result directory its config names. Inspect both bundles’ validity and metadata before comparing them. A completed command or an available size figure alone does not establish numerical correctness. hpx compare takes the two result directories:

Terminal window
hpx profile --config hpx_rt.yml
hpx profile --config hpx_aot.yml
hpx compare results/comparison_rt results/comparison_aot --output-dir results/rt-vs-aot

hpx compare reports the totals for both runs and the largest per-layer deltas, and writes compare_summary.json with the structured result. Before it compares anything it applies its own comparability rules, which the hpx compare reference states in full:

  • model identity and result validity govern whether metrics can be compared; some validity failures apply to a particular metric group;
  • a different layer topology can prevent per-layer matching;
  • power scope and capture integrity determine whether power metrics are usable;
  • engine, toolchain, board, clock, transport and placement differences are recorded comparison dimensions.

Read the comparison’s omitted metrics and warnings as part of the result. A missing delta is not a zero delta, and totals being comparable does not make every layer or power measurement comparable.

Keep the controls you can hold equal aligned between the two configs. Record remaining differences, including the runtime’s release-with-logs variant and the generated AOT module’s build settings. An engine comparison with such differences does not isolate compilation as the sole cause of a change.

Use the clean inference-window metric for whole-model timing. Per-layer counters come from an instrumented run and can have a different total. Binary figures describe the complete profiling image, including its harness; GNU size’s text metric includes code and read-only data. Memory categories can move in opposite directions, so compare them separately.

See the recorded KWS comparison for an example with model hashes, build modes and section sizes. For power measurements, follow the profiler power guide and quote the board rail and interval measured.

Keep a candidate only after its numerical validation and the metrics relevant to your application’s budget pass. Retain both original bundles and compare_summary.json so a later reader can inspect model identity, omitted metrics and build differences. Repeating a run with changed clock, placement or instrumentation is a new experiment; label it accordingly.