# Performance

Use these mechanisms to choose what to measure on your model and board. Their effect depends on the generated code, kernels and placement. Published figures are on the
[MLPerf Tiny benchmark](https://ambiqai.github.io/helia-aot/guide/benchmarks/) page, with the versions
and available measurement or report dates, and
[Measure a conversion](https://ambiqai.github.io/helia-aot/guide/measure/) is how you take your own.

## Establish a baseline

Start with a module that passes [numerical validation](https://ambiqai.github.io/helia-aot/guide/testing/)
with an appropriate oracle and tolerance. Keep its model, resolved
configuration, residency report, compiler settings and firmware image identity.
Measure it on the target and record the clock and measurement settings.

Then change one control, reconvert to a separate output directory, rebuild and
validate again. Compare occupied memory and whole-inference timing before
keeping the change. A conversion that emits fewer bytes can still produce a
slower or larger firmware image because kernel selection, linker settings and
the application harness also matter.

## Choose what to optimize

| Constraint | First check | Try next |
| --- | --- | --- |
| Firmware image size | Linker map and retained kernel code | Compare [table and static schedules](https://ambiqai.github.io/helia-aot/examples/static-schedule/) |
| Runtime memory | Arena `used` values by bank | Compare [memory planners](https://ambiqai.github.io/helia-aot/guide/memory-planners/) under the same budget |
| Inference latency | Whole-model timing and costly layers | Adjust [tensor placement](https://ambiqai.github.io/helia-aot/guide/memory-placement/) and measure again |
| Energy per inference | A valid power capture over the inference interval | Compare complete runs with the same measurement setup |

## Why the code is small

**No model interpreter.** Parsing and operator resolution happen during conversion. The generated model runs C operations: the default table schedule uses function pointers, while the static schedule emits direct calls.

**Only required operator code is emitted.** The module emits the operators and
kernel calls selected for your graph. The final firmware also links kernel
libraries and an application harness; compiler and linker settings determine
which unused library code is discarded. Use the built image to assess the
result.

**Interned constants.** Constants with identical contents are aliased onto one
copy before anything is placed, so a weight repeated across layers is stored
once.

**Interned kernel bodies.** An operator that declares the shared-kernel
contract emits its body once into the module's kernels file; each node
contributes a small read-only descriptor and a thin wrapper instead of another
copy of the body.

**Elided initialization.** An operator that declares itself stateless emits no
per-node init function at all.

## Why the run is predictable

**Arenas are planned at compile time.** Tensor lifetimes, placements and
offsets are decided during conversion. The generated module does not need to plan tensor storage at runtime; the header carries the size and alignment of every
region as a macro.

**Kernels are chosen at compile time.** Each node's kernel was selected against
the shapes and the target's capabilities during
[Resolve](https://ambiqai.github.io/helia-aot/guide/how-it-works/#resolve). Kernel selection is not repeated at each call. The table schedule still dispatches through function pointers; the static schedule uses direct calls.

**Placement is yours.** Where weights, activations and operator scratch live is
a configuration decision, not a heuristic applied on the device. See
[Memory](https://ambiqai.github.io/helia-aot/guide/memory/).

## The settings that move the numbers

Keep conversion settings in a file so you can review and repeat each experiment. Precision is determined by the model tensors, while the settings below control scheduling, kernel eligibility and placement.

| Lever | Config key | What it does | Cost |
| --- | --- | --- | --- |
| Schedule mode | `module.schedule: static` | Straight-line direct calls instead of a function-pointer loop, so the compiler sees the whole sequence. Drops the per-node callback seam. | No per-node callback, so no progress hook and no per-node instrumentation. Slightly more code for a long graph. |
| Kernel selection | `platform.name`; a complete custom target when needed | Kernel eligibility depends on tensor shapes, dtypes, operator options and target capabilities. Capability fields are ignored for registered targets. | A custom capability declaration must match the silicon; it cannot make unsupported instructions run. |
| Placement, tensors | `memory.tensors[].attributes.memory` | Moves hot activations and weights into fast memory, and large cold data out of it. | Tightly coupled memory is small, and it is the same budget your application wants. |
| Placement, constants | `memory.tensors[].attributes.constant_destination_memory` | Stages weights a kernel reads repeatedly into a memory with a lower per-access cost. | A writable runtime copy of those weights, plus a one-time transfer at init. |
| FP16 packed weights | `optimization.accumulation: fast` (or `goal: latency` with `allow_approximate: true`) | Packs FP16 1x1 convolution and fully connected weights so output channels become vector lanes (3.22x on an Apollo510 pointwise layer). FP32 packs by default. See [packed float weights](https://ambiqai.github.io/helia-aot/guide/precision/#packed-float-weights). | FP16-lane accumulation: roughly 4-5x the standard layout's error from 32 input channels up. |
| Code placement | `operators[].attributes.code_placement` | Requests ITCM for generated operator run code; kernel-library placement follows the linker script. See [ITCM placement](https://ambiqai.github.io/helia-aot/guide/itcm-placement/). | Code and explicitly placed tensors share the ITCM budget. |
| Precision | The quantization of the model you convert, and `platform.name` for what the target can run | Quantization can reduce storage; speed depends on the model, operator and target kernels. FP16/FP32 support is experimental and operator-specific. Native FP16 computation requires compatible target support; FP16 weight storage can use software widening. | Accuracy requirements and operator-specific code, constant and workspace costs; compare the complete footprint. |

Where each one is documented in full: [Schedule
mode](https://ambiqai.github.io/helia-aot/guide/configuring/#schedule-mode), [kernel
selection](https://ambiqai.github.io/helia-aot/guide/how-it-works/#kernel-selection),
[Memory](https://ambiqai.github.io/helia-aot/guide/memory/#choosing-where-a-tensor-goes), [cold and
staged constants](https://ambiqai.github.io/helia-aot/guide/memory/#cold-and-staged-constants),
[operator attributes](https://ambiqai.github.io/helia-aot/guide/configuring/#operator-attributes) and
[Precision](https://ambiqai.github.io/helia-aot/guide/precision/).

There is no configuration switch that converts an integer model into an FP16
or FP32 model. Precision comes from its tensors, and each operator enforces its
own dtype, shape and option restrictions. For example, native rolled GRU uses
FP16 throughout; that does not imply every operator supports FP16. FP16 weight
storage is a separate case: `DEQUANTIZE` can widen it to FP32 in software on a
target without native FP16 support. Check the
[operator catalog](https://ambiqai.github.io/helia-aot/reference/operators/) and the required
[float library switches](https://ambiqai.github.io/helia-aot/guide/precision/#the-library-build-switch).

The options that trade speed against accuracy or memory, with their measured
gains and costs, are listed on
[Performance and accuracy options](https://ambiqai.github.io/helia-aot/guide/options/); each conversion
writes its choice for every operator to `<prefix>_plan.json` and prints the
alternatives that apply to it as optimization hints.

Change one control at a time and measure between changes. The levers interact:
staged weights and activations compete for the same runtime memory budget.

Tightly coupled memory is scarce. Promote the few operators that dominate the
run, measure, and put the rest back.

## Reading a profiler run

heliaPROFILER (`hpx`) builds firmware around your model, runs it on the board,
and writes a result bundle. Read it in this order:

1. **Did it run, and is it right?** A conversion that compiles has not been
   validated. The generated test case can check supplied golden outputs or host-interpreter
   outputs within `test.tolerance`; keep `skip_verification` false. Numbers from a run
   that does not match the reference are not worth comparing.
2. **Whole-inference timing**, measured in the clean inference window and interpreted with the recorded processor clock. This is
   the figure that answers whether the model fits the time budget.
3. **The per-layer rows.** They are where the time actually is. Expect it to be
   concentrated: a handful of nodes usually dominate, and they are the only
   ones worth moving into fast memory or re-quantizing.
4. **The counters.** Available groups depend on the processor and selected profiling backend. Use stall, memory and vector counters as evidence to investigate a slow layer; a single counter does not establish its cause.
5. **Binary size.** Reported for the firmware the profiler built, so it is
   comparable between two runs of the same harness and not with a figure from
   a different build.
6. **Energy**, when a supported power monitor and the matching wiring and measurement mode are configured. Read the measurement interval and capture-health information before comparing energy figures.

`hpx compare` takes two bundles and reports the deltas. Because the engine is
one of the things a run varies, the same model on the same board can be
measured through heliaAOT, heliaRT or TFLM. Compare the recorded build and
measurement conditions along with the results. The recipe is
on [Measure a conversion](https://ambiqai.github.io/helia-aot/guide/measure/).

:::note
Compare bundles from the same board, the same toolchain and the same counter
selection. A delta across any of those is measuring the difference in the
harness as much as the difference in the module.
:::
