Skip to content
heliaAOT
HELIA HUB

Performance

Use these mechanisms to choose what to measure on your model and board. Their effect depends on the generated code, kernels and placement. Published figures are on the MLPerf Tiny benchmark page, with the versions and available measurement or report dates, and Measure a conversion is how you take your own.

Start with a module that passes numerical validation with an appropriate oracle and tolerance. Keep its model, resolved configuration, residency report, compiler settings and firmware image identity. Measure it on the target and record the clock and measurement settings.

Then change one control, reconvert to a separate output directory, rebuild and validate again. Compare occupied memory and whole-inference timing before keeping the change. A conversion that emits fewer bytes can still produce a slower or larger firmware image because kernel selection, linker settings and the application harness also matter.

Constraint First check Try next
Firmware image size Linker map and retained kernel code Compare table and static schedules
Runtime memory Arena used values by bank Compare memory planners under the same budget
Inference latency Whole-model timing and costly layers Adjust tensor placement and measure again
Energy per inference A valid power capture over the inference interval Compare complete runs with the same measurement setup

No model interpreter. Parsing and operator resolution happen during conversion. The generated model runs C operations: the default table schedule uses function pointers, while the static schedule emits direct calls.

Only required operator code is emitted. The module emits the operators and kernel calls selected for your graph. The final firmware also links kernel libraries and an application harness; compiler and linker settings determine which unused library code is discarded. Use the built image to assess the result.

Interned constants. Constants with identical contents are aliased onto one copy before anything is placed, so a weight repeated across layers is stored once.

Interned kernel bodies. An operator that declares the shared-kernel contract emits its body once into the module’s kernels file; each node contributes a small read-only descriptor and a thin wrapper instead of another copy of the body.

Elided initialization. An operator that declares itself stateless emits no per-node init function at all.

Arenas are planned at compile time. Tensor lifetimes, placements and offsets are decided during conversion. The generated module does not need to plan tensor storage at runtime; the header carries the size and alignment of every region as a macro.

Kernels are chosen at compile time. Each node’s kernel was selected against the shapes and the target’s capabilities during Resolve. Kernel selection is not repeated at each call. The table schedule still dispatches through function pointers; the static schedule uses direct calls.

Placement is yours. Where weights, activations and operator scratch live is a configuration decision, not a heuristic applied on the device. See Memory.

Keep conversion settings in a file so you can review and repeat each experiment. Precision is determined by the model tensors, while the settings below control scheduling, kernel eligibility and placement.

Lever Config key What it does Cost
Schedule mode module.schedule: static Straight-line direct calls instead of a function-pointer loop, so the compiler sees the whole sequence. Drops the per-node callback seam. No per-node callback, so no progress hook and no per-node instrumentation. Slightly more code for a long graph.
Kernel selection platform.name; a complete custom target when needed Kernel eligibility depends on tensor shapes, dtypes, operator options and target capabilities. Capability fields are ignored for registered targets. A custom capability declaration must match the silicon; it cannot make unsupported instructions run.
Placement, tensors memory.tensors[].attributes.memory Moves hot activations and weights into fast memory, and large cold data out of it. Tightly coupled memory is small, and it is the same budget your application wants.
Placement, constants memory.tensors[].attributes.constant_destination_memory Stages weights a kernel reads repeatedly into a memory with a lower per-access cost. A writable runtime copy of those weights, plus a one-time transfer at init.
FP16 packed weights optimization.accumulation: fast (or goal: latency with allow_approximate: true) Packs FP16 1x1 convolution and fully connected weights so output channels become vector lanes (3.22x on an Apollo510 pointwise layer). FP32 packs by default. See packed float weights. FP16-lane accumulation: roughly 4-5x the standard layout’s error from 32 input channels up.
Code placement operators[].attributes.code_placement Requests ITCM for generated operator run code; kernel-library placement follows the linker script. See ITCM placement. Code and explicitly placed tensors share the ITCM budget.
Precision The quantization of the model you convert, and platform.name for what the target can run Quantization can reduce storage; speed depends on the model, operator and target kernels. FP16/FP32 support is experimental and operator-specific. Native FP16 computation requires compatible target support; FP16 weight storage can use software widening. Accuracy requirements and operator-specific code, constant and workspace costs; compare the complete footprint.

Where each one is documented in full: Schedule mode, kernel selection, Memory, cold and staged constants, operator attributes and Precision.

There is no configuration switch that converts an integer model into an FP16 or FP32 model. Precision comes from its tensors, and each operator enforces its own dtype, shape and option restrictions. For example, native rolled GRU uses FP16 throughout; that does not imply every operator supports FP16. FP16 weight storage is a separate case: DEQUANTIZE can widen it to FP32 in software on a target without native FP16 support. Check the operator catalog and the required float library switches.

The options that trade speed against accuracy or memory, with their measured gains and costs, are listed on Performance and accuracy options; each conversion writes its choice for every operator to <prefix>_plan.json and prints the alternatives that apply to it as optimization hints.

Change one control at a time and measure between changes. The levers interact: staged weights and activations compete for the same runtime memory budget.

Tightly coupled memory is scarce. Promote the few operators that dominate the run, measure, and put the rest back.

heliaPROFILER (hpx) builds firmware around your model, runs it on the board, and writes a result bundle. Read it in this order:

  1. Did it run, and is it right? A conversion that compiles has not been validated. The generated test case can check supplied golden outputs or host-interpreter outputs within test.tolerance; keep skip_verification false. Numbers from a run that does not match the reference are not worth comparing.
  2. Whole-inference timing, measured in the clean inference window and interpreted with the recorded processor clock. This is the figure that answers whether the model fits the time budget.
  3. The per-layer rows. They are where the time actually is. Expect it to be concentrated: a handful of nodes usually dominate, and they are the only ones worth moving into fast memory or re-quantizing.
  4. The counters. Available groups depend on the processor and selected profiling backend. Use stall, memory and vector counters as evidence to investigate a slow layer; a single counter does not establish its cause.
  5. Binary size. Reported for the firmware the profiler built, so it is comparable between two runs of the same harness and not with a figure from a different build.
  6. Energy, when a supported power monitor and the matching wiring and measurement mode are configured. Read the measurement interval and capture-health information before comparing energy figures.

hpx compare takes two bundles and reports the deltas. Because the engine is one of the things a run varies, the same model on the same board can be measured through heliaAOT, heliaRT or TFLM. Compare the recorded build and measurement conditions along with the results. The recipe is on Measure a conversion.