Skip to content
heliaAOT
HELIA HUB

Performance and accuracy options

heliaAOT’s defaults keep the kernel library’s default numerics. The optimization settings let a conversion trade accuracy or memory for speed where your model allows it. Every conversion writes the choice it made for each operator to <prefix>_plan.json, explains it in <prefix>_report.json, and prints the alternatives that apply under Optimization hints in the results, with the operators they would affect, their share of the model’s MACs, the measured gain and the line of YAML that sets them. Hints are facts only; they never change the conversion.

  • Setting: anything below that you set: the goal, the approximation gate, each knob, and constant placement.
  • Knob: a per-operator mechanism with a closed set of values, plus auto.
  • Exact and approximate: every knob value is classified. An approximate value changes numerics relative to the default value of its knob; auto picks one only when allow_approximate is true.
  • Goal: what auto favors: latency, balanced, size or accuracy.
  • Plan: the fully resolved choice of every knob for every operator, written to <prefix>_plan.json.
  • Report: the facts behind the plan and its hints, written to <prefix>_report.json.
  • auto: a static, documented table in heliaAOT that resolves a knob from the goal. It never measures or searches; searching for the best plan on a target happens outside the compiler.
Setting Set with Default Values Measured Accuracy cost
Goal optimization.goal balanced latency, balanced, size, accuracy - -
Approximation gate optimization.allow_approximate false false, true - -
Accumulation optimization.accumulation, operators[].optimization.accumulation auto auto, precise (the default accumulation), fast (approximate) fast: 3.22x on an FP16 pointwise layer (1x256, 32 to 32 channels), Apollo510 fast: max |error| 0.0156 against 0.0059 with precise on that layer
Kernel optimization.kernel, operators[].optimization.kernel auto auto, specialized (shape-specialized direct entries where they apply), generic (the general entries) ns-cmsis-nn 7.37.0 kernel-level, Apollo510: 1.9-2.4x on first-layer convolutions with 1-3 input channels, 1.23x on 3x3 convolutions with 16 channels, 1.7-2.0x on 3x3 depthwise layers Exact: identical output
Constant placement memory.tensors[].attributes.constant_destination_memory unset: constants are read where the planner places them a memory, e.g. DTCM DTCM instead of SRAM, whole model on Apollo510 at 96 MHz: AD 2.70x, sleep FP16 1.28x, sleep FP32 1.22x; KWS, VWW, IC and TCN within 1% Exact; costs a DTCM copy of the weights and an init-time copy

The accumulation and constant placement measurements used heliaAOT 0.23.0 and ns-cmsis-nn 7.36.0 on an Apollo510 EVB (Cortex-M55). Each is a record in heliaAOT’s measurement data, with its target, core, versions and source. The report carries only the records measured on the conversion’s target and core; on any other target, even one with the same core, the hints state the facts without a measured gain. The kernel figures are ns-cmsis-nn’s own kernel-level measurements, not heliaAOT conversions, so they are not in the measurement data and no report quotes them yet.

optimization.budget and optimization.plan are reserved for an optimizer that searches plans on a target and for reproducing a searched plan; setting either is an error today. Further knobs (weight layout, activation LUTs, code specialization, placement) will join the same section.

Knob latency with allow_approximate: true latency balanced size accuracy
accumulation fast precise precise precise precise
kernel specialized specialized specialized generic specialized

auto never picks an approximate value unless allow_approximate is true. A value you set explicitly, for the model or an operator, is used as given: setting accumulation: fast is your consent to its numerics, and the plan records it as approximate.

accumulation applies to FP16 1x1 convolutions and fully connected layers. With precise, the weights keep the standard layout and the kernels use ns-cmsis-nn’s default accumulation: today FP16 partial sums in vector lanes on MVE, and blockwise FP16 partials folded into FP32 once the library provides it. With fast, the weights are packed so output channels become vector lanes and each output accumulates in one long FP16 chain, which is approximate relative to precise. On MVE both accumulate in FP16 today; the difference is how many partial sums the kernel keeps. Packing pays off only on an MVE FP16 target with at least 4 rows sharing the weights and at least one full lane group of 8 outputs; other layers keep the standard layout, and the report marks the knob not applicable there with the reason. See packed float weights for the layout and the accuracy analysis. FP32 weights pack where it pays off regardless of this knob, and keep FP32 accumulation.

Let auto choose for latency:

optimization:
goal: latency
allow_approximate: true

Or set the knob for the whole model, and override it per operator type or id:

optimization:
accumulation: fast
operators:
- type: FULLY_CONNECTED
id: "12"
optimization:
accumulation: precise

accumulation replaces the fp16_accuracy operator attribute of an earlier development build; that name was never released, so it has no alias.

kernel applies to int8 convolutions and depthwise convolutions on MVE targets; other targets keep today’s kernels. With specialized, a layer calls the ns-cmsis-nn 7.37.0 direct entry whose shape rule it meets:

  • arm_convolve_s8_small_cin for 1 to 3 input channels;
  • arm_convolve_s8_3x3_c16_s1 for 3x3 kernels over 16 channels with stride 1;
  • arm_depthwise_conv_s8_opt_3x3_c64_s1 and arm_depthwise_conv_s8_opt_3x3 for 3x3 depthwise layers;
  • arm_depthwise_conv_s8_opt_planar where the planar rule holds;
  • arm_depthwise_conv_s8_opt_channelwise for every other layer the optimized depthwise kernel takes.

With generic, no layer calls a 7.37.0 direct entry: a layer that would take one calls arm_convolve_s8 or arm_depthwise_conv_s8_opt instead. The existing shape-specific kernels (1x1, 1xN, and arm_depthwise_conv_3x3_s8 on DSP targets) apply with either value. Both values give identical output. Each layer calls one entry directly, with no run-time dispatch.

generic is not always the smaller module. It avoids linking several direct entries next to the general ones, but arm_depthwise_conv_s8_opt carries both its planar and channel paths, and a 3x3 depthwise layer needs a larger scratch buffer with it. Code size with specialized against generic, linked for the Cortex-M55 with unused sections removed:

  • the MLPerf Tiny keyword-spotting model is 1928 B smaller, with 1120 B less arena;
  • the visual wake words model is 6520 B larger;
  • the image classification model is 2652 B larger.

goal: size picks generic; measure your model if the difference matters. A module that calls a direct entry requires ns-cmsis-nn 7.37.0 or later; other modules keep their floor. The report names the entry every operator calls under entry.

Constants are read-only, so their cold copy stays in non-volatile memory, set with memory. constant_destination_memory adds a runtime copy in a faster memory, which the generated <prefix>_model_init fills once. Staging weights into DTCM removes the SRAM weight-read stalls of weight-bound layers such as fully connected and conv1d; activation-bound models change little. DTCM is shared with your application’s stack and data, so the report states how much of it the module leaves free, within memory.constraints. See memory placement.

memory:
tensors:
- type: constant
attributes: {memory: MRAM, constant_destination_memory: DTCM}

<prefix>_plan.json sits next to the module and holds the choices. It carries a schema_version, the heliaAOT version and the ns-cmsis-nn version the module requires, the model it was made from (file name, sha256 and subgraph), the target (name, core, capabilities, FP16 support), the goal and approximation gate, and one entry per operator under operators.

Each operator has a stable id, "<subgraph>:<output tensor name>", taken from the name the model gives the operator’s first output tensor, so it stays the same when other operators are added or removed. Its position in execution order is kept in index, and its AIR node id in node_id. Repeated names get #1, #2 in operator order, and an operator without a name falls back to "<subgraph>:<op_type>@<position>"; id_source says which applied. A tensor name can itself contain : (for example 0:StatefulPartitionedCall_1:0), so a parser splits these ids on their first colon only. An id with id_source: node_id (see below) is the opaque AIR node id and is not split. operators[].id rules match the stable id as well as the node id, so a plan’s ids can be written back as rules. A stable id that would also select another operator (another operator’s node id, or a stable id spelled twice) is not matched; the plan then uses the node id as id, with id_source: node_id, and the conversion warns.

For each knob of an operator the plan records the value used, the value requested, resolved_by (explicit, or auto_table when auto chose it from the goal; auto_heuristic and plan are reserved), whether the value is approximate where it applies, and the source: kind is operator (with rule, the position, type and id of the winning operators entry), model (set under optimization, even to auto), or default (plan is reserved). Under fixed, it records choices that change numerics or speed but are not knobs yet, with what set them (attribute, default, heuristic or knob; plan is reserved): RSQRT lut_mode (INT16 only), HARD_SWISH use_lut and compat_variant (INT8), and the float weight layout of 1x1 convolutions and fully connected layers (weight_layout: set by accumulation for FP16, except where fast leaves the standard layout because packing would not pay off, and by a heuristic otherwise).

<prefix>_report.json explains the plan and does not change it: each operator’s MACs, the alternatives of each knob and whether they apply on this target (and why not), the measurements for this target and core keyed by setting, the constant placement facts, and the hints the results print. Its content may change between releases; the choices are in the plan.

Both files carry a schema_version. Adding an optional field, a knob, a knob value or an enum member keeps it; renaming or removing a field, or changing what one means, increments it. The JSON schemas are exported with the documentation data (plan-schema.json, report-schema.json) and list every knob with its values, so they belong to the heliaAOT release that wrote a file (generator.helia_aot): validate a file against that release’s schema rather than pinning one. helia_aot.optimization reads both files back as typed models. Some values and fields are reserved and empty today: plan as a source, optimization.knobs for model-level choices and tensors for per-tensor choices.

Knobs are set only under optimization, for the model or in an operators rule; a knob name under attributes is rejected. optimization.budget and optimization.plan are YAML-only while they are reserved.

A hint appears only when its alternative applies and a single setting would select it. A knob hint lists the operators where the knob applies that do not run the value auto picks for goal: latency with allow_approximate: true (for accumulation, operators running precise that fast would pack), leaving out operators an operators rule pins. The placement hint needs constants outside TCM that all come from one cold memory and fit in the free DTCM. Share of MACs counts convolution, depthwise and fully connected layers; it is not a share of run time unless you profile the model.