Measure a conversion
A conversion is worth measuring on the Ambiq silicon it will run on. heliaPROFILER
(hpx) builds profiler firmware around a model, flashes a target, captures
what the run cost, and writes a result directory you can compare against
another one.
Before you measure
Section titled “Before you measure”First validate the module with a trusted oracle and appropriate tolerance. You need a profiler-supported board, toolchain and probe configured for that board. For power, complete the profiler’s monitor and wiring setup as well. These are separate prerequisites from successfully converting the model.
Keep the model file and its hash, target, clock, engine versions, compiler settings and placement with the result. The commands below operate hardware; a successful conversion alone does not produce a measurement bundle.
What hpx measures
Section titled “What hpx measures”hpx profile measures whole-inference timing, optional per-layer measurements and the firmware image it builds. Available counters depend on the processor. Power and energy require a supported monitor, correct wiring and the corresponding measurement mode. Every option is in the
hpx profile reference.
The engine is one of the things it varies, so the same model can be measured
through heliaRT, heliaAOT or TFLM. Keep the model, board, clock, toolchain and measurement settings aligned, then inspect the differences recorded in the bundles.
Install it
Section titled “Install it”pip install 'helia-profiler[aot]'hpx doctorhpx doctor checks installed host tools and Python dependencies. It does not
verify a connected probe or board, and its default checks do not establish AOT
readiness. Install the aot extra and follow the board setup instructions. See the
heliaPROFILER documentation for
the full setup. heliaPROFILER has its own release and configuration contract;
record the installed version and follow that version’s reference if it differs
from the recipe below. An AOT package version does not pin the profiler,
runtime engine or toolchain for you.
One config per engine
Section titled “One config per engine”The recipe below is heliaPROFILER’s
engine comparison example,
adapted here with matching model, board and profiling settings. Replace my_model.tflite with your model path. The engine configuration and output directory differ; generated code and runtime build modes can differ too.
model: path: my_model.tflite arena_size: 131072
engine: type: helia-rt config: variant: release-with-logs
target: board: apollo510_evb toolchain: arm-none-eabi-gcc
profiling: pmu_counters: cpu: default per_layer: true iterations: 5 warmup: 2
output: format: csv dir: ./results/comparison_rt detailed: truemodel: path: my_model.tflite arena_size: 131072
engine: type: helia-aot config: prefix: hpx module_name: hpx_model
target: board: apollo510_evb toolchain: arm-none-eabi-gcc
profiling: pmu_counters: cpu: default per_layer: true iterations: 5 warmup: 2
output: format: csv dir: ./results/comparison_aot detailed: trueRun both, then compare
Section titled “Run both, then compare”Each successful run writes the result directory its config names. Inspect
both bundles’ validity and metadata before comparing them. A completed command
or an available size figure alone does not establish numerical correctness.
hpx compare takes the two result directories:
hpx profile --config hpx_rt.ymlhpx profile --config hpx_aot.ymlhpx compare results/comparison_rt results/comparison_aot --output-dir results/rt-vs-aotHow to read the comparison
Section titled “How to read the comparison”hpx compare reports the totals for both runs and the largest per-layer
deltas, and writes compare_summary.json with the structured result. Before
it compares anything it applies its own comparability rules, which the
hpx compare reference
states in full:
- model identity and result validity govern whether metrics can be compared; some validity failures apply to a particular metric group;
- a different layer topology can prevent per-layer matching;
- power scope and capture integrity determine whether power metrics are usable;
- engine, toolchain, board, clock, transport and placement differences are recorded comparison dimensions.
Read the comparison’s omitted metrics and warnings as part of the result. A missing delta is not a zero delta, and totals being comparable does not make every layer or power measurement comparable.
Keep the controls you can hold equal aligned between the two configs. Record
remaining differences, including the runtime’s release-with-logs variant
and the generated AOT module’s build settings. An engine comparison with such
differences does not isolate compilation as the sole cause of a change.
Use the clean inference-window metric for whole-model timing. Per-layer counters come from an instrumented run and can have a different total. Binary figures describe the complete profiling image, including its harness; GNU size’s text metric includes code and read-only data. Memory categories can move in opposite directions, so compare them separately.
See the recorded KWS comparison for an example with model hashes, build modes and section sizes. For power measurements, follow the profiler power guide and quote the board rail and interval measured.
Keep or reject a candidate
Section titled “Keep or reject a candidate”Keep a candidate only after its numerical validation and the metrics relevant
to your application’s budget pass. Retain both original bundles and
compare_summary.json so a later reader can inspect model identity, omitted
metrics and build differences. Repeating a run with changed clock, placement
or instrumentation is a new experiment; label it accordingly.
Where to go deeper
Section titled “Where to go deeper”- Analysis and run comparison
for
hpx analyzeand the comparison workflow. - Inference engines for what each engine is and what it costs.
- MLPerf Tiny benchmark for a comparison already run on this board.