Performance and accuracy options
heliaAOT’s defaults keep the kernel library’s default numerics. The optimization settings let a conversion
trade accuracy or memory for speed where your model allows it. Every conversion
writes the choice it made for each operator to <prefix>_plan.json, explains
it in <prefix>_report.json, and prints the alternatives that apply under
Optimization hints in the results, with the operators they would affect, their share of the model’s MACs, the
measured gain and the line of YAML that sets them. Hints are facts only; they
never change the conversion.
Vocabulary
Section titled “Vocabulary”- Setting: anything below that you set: the goal, the approximation gate, each knob, and constant placement.
- Knob: a per-operator mechanism with a closed set of values, plus
auto. - Exact and approximate: every knob value is classified. An approximate
value changes numerics relative to the default value of its knob;
autopicks one only whenallow_approximateistrue. - Goal: what
autofavors:latency,balanced,sizeoraccuracy. - Plan: the fully resolved choice of every knob for every operator, written
to
<prefix>_plan.json. - Report: the facts behind the plan and its hints, written to
<prefix>_report.json. - auto: a static, documented table in heliaAOT that resolves a knob from the goal. It never measures or searches; searching for the best plan on a target happens outside the compiler.
The settings
Section titled “The settings”| Setting | Set with | Default | Values | Measured | Accuracy cost |
|---|---|---|---|---|---|
| Goal | optimization.goal |
balanced |
latency, balanced, size, accuracy |
- | - |
| Approximation gate | optimization.allow_approximate |
false |
false, true |
- | - |
| Accumulation | optimization.accumulation, operators[].optimization.accumulation |
auto |
auto, precise (the default accumulation), fast (approximate) |
fast: 3.22x on an FP16 pointwise layer (1x256, 32 to 32 channels), Apollo510 |
fast: max |error| 0.0156 against 0.0059 with precise on that layer |
| Kernel | optimization.kernel, operators[].optimization.kernel |
auto |
auto, specialized (shape-specialized direct entries where they apply), generic (the general entries) |
ns-cmsis-nn 7.37.0 kernel-level, Apollo510: 1.9-2.4x on first-layer convolutions with 1-3 input channels, 1.23x on 3x3 convolutions with 16 channels, 1.7-2.0x on 3x3 depthwise layers | Exact: identical output |
| Constant placement | memory.tensors[].attributes.constant_destination_memory |
unset: constants are read where the planner places them |
a memory, e.g. DTCM |
DTCM instead of SRAM, whole model on Apollo510 at 96 MHz: AD 2.70x, sleep FP16 1.28x, sleep FP32 1.22x; KWS, VWW, IC and TCN within 1% |
Exact; costs a DTCM copy of the weights and an init-time copy |
The accumulation and constant placement measurements used heliaAOT 0.23.0 and ns-cmsis-nn 7.36.0 on an Apollo510 EVB (Cortex-M55). Each is a record in heliaAOT’s measurement data, with its target, core, versions and source. The report carries only the records measured on the conversion’s target and core; on any other target, even one with the same core, the hints state the facts without a measured gain. The kernel figures are ns-cmsis-nn’s own kernel-level measurements, not heliaAOT conversions, so they are not in the measurement data and no report quotes them yet.
optimization.budget and optimization.plan are reserved for an optimizer
that searches plans on a target and for reproducing a searched plan; setting
either is an error today. Further knobs (weight layout, activation LUTs, code
specialization, placement) will join the same section.
How auto resolves
Section titled “How auto resolves”| Knob | latency with allow_approximate: true |
latency |
balanced |
size |
accuracy |
|---|---|---|---|---|---|
accumulation |
fast |
precise |
precise |
precise |
precise |
kernel |
specialized |
specialized |
specialized |
generic |
specialized |
auto never picks an approximate value unless allow_approximate is true. A value
you set explicitly, for the model or an operator, is used as given: setting
accumulation: fast is your consent to its numerics, and the plan records it
as approximate.
Accumulation
Section titled “Accumulation”accumulation applies to FP16 1x1 convolutions and fully connected layers.
With precise, the weights keep the standard layout and the kernels use
ns-cmsis-nn’s default accumulation: today FP16 partial sums in vector lanes on
MVE, and blockwise FP16 partials folded into FP32 once the library provides
it. With fast, the weights are packed so output channels become vector lanes
and each output accumulates in one long FP16 chain, which is approximate
relative to precise. On MVE both accumulate in FP16 today; the difference is
how many partial sums the kernel keeps. Packing pays off only on an MVE
FP16 target with at least 4 rows sharing the weights and at least one full lane
group of 8 outputs; other layers keep the standard layout, and the report marks
the knob not applicable there with the reason. See
packed float weights for
the layout and the accuracy analysis. FP32 weights pack where it pays off
regardless of this knob, and keep FP32 accumulation.
Let auto choose for latency:
optimization: goal: latency allow_approximate: trueOr set the knob for the whole model, and override it per operator type or id:
optimization: accumulation: fastoperators: - type: FULLY_CONNECTED id: "12" optimization: accumulation: preciseaccumulation replaces the fp16_accuracy operator attribute of an earlier
development build; that name was never released, so it has no alias.
Kernel
Section titled “Kernel”kernel applies to int8 convolutions and depthwise convolutions on MVE
targets; other targets keep today’s kernels. With specialized, a layer calls
the ns-cmsis-nn 7.37.0 direct entry whose shape rule it meets:
arm_convolve_s8_small_cinfor 1 to 3 input channels;arm_convolve_s8_3x3_c16_s1for 3x3 kernels over 16 channels with stride 1;arm_depthwise_conv_s8_opt_3x3_c64_s1andarm_depthwise_conv_s8_opt_3x3for 3x3 depthwise layers;arm_depthwise_conv_s8_opt_planarwhere the planar rule holds;arm_depthwise_conv_s8_opt_channelwisefor every other layer the optimized depthwise kernel takes.
With generic, no layer calls a 7.37.0 direct entry: a layer that would take
one calls arm_convolve_s8 or arm_depthwise_conv_s8_opt instead. The
existing shape-specific kernels (1x1, 1xN, and arm_depthwise_conv_3x3_s8 on
DSP targets) apply with either value. Both values give identical output. Each
layer calls one entry directly, with no run-time dispatch.
generic is not always the smaller module. It avoids linking several direct
entries next to the general ones, but arm_depthwise_conv_s8_opt carries both
its planar and channel paths, and a 3x3 depthwise layer needs a larger scratch
buffer with it. Code size with specialized against generic, linked for the
Cortex-M55 with unused sections removed:
- the MLPerf Tiny keyword-spotting model is 1928 B smaller, with 1120 B less arena;
- the visual wake words model is 6520 B larger;
- the image classification model is 2652 B larger.
goal: size picks generic; measure your model if the difference matters. A module that calls
a direct entry requires ns-cmsis-nn 7.37.0 or later; other modules keep their
floor. The report names the entry every operator calls under entry.
Constant placement
Section titled “Constant placement”Constants are read-only, so their cold copy stays in non-volatile memory, set
with memory. constant_destination_memory adds a runtime copy in a faster
memory, which the generated <prefix>_model_init fills once. Staging weights
into DTCM removes the SRAM weight-read stalls of weight-bound layers such as
fully connected and conv1d; activation-bound models change little. DTCM is
shared with your application’s stack and data, so the report states how much
of it the module leaves free, within memory.constraints. See
memory placement.
memory: tensors: - type: constant attributes: {memory: MRAM, constant_destination_memory: DTCM}The plan and the report
Section titled “The plan and the report”<prefix>_plan.json sits next to the module and holds the choices. It carries
a schema_version, the heliaAOT version and the ns-cmsis-nn version the module
requires, the model it was made from (file name, sha256 and subgraph), the
target (name, core, capabilities, FP16 support), the goal and approximation
gate, and one entry per operator under operators.
Each operator has a stable id, "<subgraph>:<output tensor name>", taken
from the name the model gives the operator’s first output tensor, so it stays
the same when other operators are added or removed. Its position in execution
order is kept in index, and its AIR node id in node_id. Repeated names get
#1, #2 in operator order, and an operator without a name falls back to
"<subgraph>:<op_type>@<position>"; id_source says which applied. A tensor
name can itself contain : (for example 0:StatefulPartitionedCall_1:0), so
a parser splits these ids on their first colon only. An id with
id_source: node_id (see below) is the opaque AIR node id and is not split.
operators[].id rules match the stable id as well as the node id, so a plan’s
ids can be written back as rules. A stable id that would also select another
operator (another operator’s node id, or a stable id spelled twice) is not
matched; the plan then uses the node id as id, with id_source: node_id,
and the conversion warns.
For each knob of an operator the plan records the value used, the value
requested, resolved_by (explicit, or auto_table when auto chose it
from the goal; auto_heuristic and plan are reserved), whether the value is
approximate where it applies, and the source: kind is operator (with
rule, the position, type and id of the winning operators entry), model
(set under optimization, even to auto), or default (plan is reserved).
Under fixed, it records choices that change numerics or speed but are not
knobs yet, with what set them (attribute, default, heuristic or knob;
plan is reserved): RSQRT lut_mode (INT16 only),
HARD_SWISH use_lut and compat_variant (INT8), and the float weight
layout of 1x1 convolutions and fully connected layers (weight_layout: set
by accumulation for FP16, except where fast leaves the standard layout
because packing would not pay off, and by a heuristic otherwise).
<prefix>_report.json explains the plan and does not change it: each
operator’s MACs, the alternatives of each knob and whether they apply on this
target (and why not), the measurements for this target and core keyed by
setting, the constant placement facts, and the hints the results print. Its
content may change between releases; the choices are in the plan.
Both files carry a schema_version. Adding an optional field, a knob, a knob
value or an enum member keeps it; renaming or removing a field, or changing
what one means, increments it. The JSON schemas are exported with the
documentation data (plan-schema.json, report-schema.json) and list every
knob with its values, so they belong to the heliaAOT release that wrote a file
(generator.helia_aot): validate a file against that release’s schema rather
than pinning one. helia_aot.optimization reads both files back as typed
models. Some values and fields are reserved and empty today: plan as a
source, optimization.knobs for model-level choices and tensors for
per-tensor choices.
Knobs are set only under optimization, for the model or in an operators
rule; a knob name under attributes is rejected. optimization.budget and
optimization.plan are YAML-only while they are reserved.
A hint appears only when its alternative applies and a single setting would
select it. A knob hint lists the operators where the knob applies that do not
run the value auto picks for goal: latency with allow_approximate: true
(for accumulation, operators running precise that fast would pack),
leaving out operators an operators rule pins. The placement hint needs constants outside TCM that all come from one
cold memory and fit in the free DTCM. Share of MACs counts convolution,
depthwise and fully connected layers; it is not a share of run time unless you
profile the model.