# Performance and accuracy options

heliaAOT's defaults keep the kernel library's default numerics. The `optimization` settings let a conversion
trade accuracy or memory for speed where your model allows it. Every conversion
writes the choice it made for each operator to `<prefix>_plan.json`, explains
it in `<prefix>_report.json`, and prints the alternatives that apply under
**Optimization hints** in the results, with the operators they would affect, their share of the model's MACs, the
measured gain and the line of YAML that sets them. Hints are facts only; they
never change the conversion.

## Vocabulary

- **Setting:** anything below that you set: the goal, the approximation gate,
  each knob, and constant placement.
- **Knob:** a per-operator mechanism with a closed set of values, plus `auto`.
- **Exact and approximate:** every knob value is classified. An approximate
  value changes numerics relative to the default value of its knob; `auto`
  picks one only when `allow_approximate` is `true`.
- **Goal:** what `auto` favors: `latency`, `balanced`, `size` or `accuracy`.
- **Plan:** the fully resolved choice of every knob for every operator, written
  to `<prefix>_plan.json`.
- **Report:** the facts behind the plan and its hints, written to
  `<prefix>_report.json`.
- **auto:** a static, documented table in heliaAOT that resolves a knob from the
  goal. It never measures or searches; searching for the best plan on a
  target happens outside the compiler.

## The settings

| Setting | Set with | Default | Values | Measured | Accuracy cost |
| --- | --- | --- | --- | --- | --- |
| Goal | `optimization.goal` | `balanced` | `latency`, `balanced`, `size`, `accuracy` | - | - |
| Approximation gate | `optimization.allow_approximate` | `false` | `false`, `true` | - | - |
| Accumulation | `optimization.accumulation`, `operators[].optimization.accumulation` | `auto` | `auto`, `precise` (the default accumulation), `fast` (approximate) | `fast`: 3.22x on an FP16 pointwise layer (1x256, 32 to 32 channels), Apollo510 | `fast`: max \|error\| 0.0156 against 0.0059 with `precise` on that layer |
| Kernel | `optimization.kernel`, `operators[].optimization.kernel` | `auto` | `auto`, `specialized` (shape-specialized direct entries where they apply), `generic` (the general entries) | ns-cmsis-nn 7.37.0 kernel-level, Apollo510: 1.9-2.4x on first-layer convolutions with 1-3 input channels, 1.23x on 3x3 convolutions with 16 channels, 1.7-2.0x on 3x3 depthwise layers | Exact: identical output |
| Constant placement | `memory.tensors[].attributes.constant_destination_memory` | `unset`: constants are read where the planner places them | a memory, e.g. `DTCM` | `DTCM` instead of SRAM, whole model on Apollo510 at 96 MHz: AD 2.70x, sleep FP16 1.28x, sleep FP32 1.22x; KWS, VWW, IC and TCN within 1% | Exact; costs a DTCM copy of the weights and an init-time copy |

The accumulation and constant placement measurements used heliaAOT 0.23.0 and
ns-cmsis-nn 7.36.0 on an Apollo510 EVB (Cortex-M55). Each is a record in
heliaAOT's measurement data, with its target, core, versions and source. The
report carries only the records measured on the conversion's target and core;
on any other target, even one with the same core, the hints state the facts
without a measured gain. The kernel figures are ns-cmsis-nn's own kernel-level
measurements, not heliaAOT conversions, so they are not in the measurement
data and no report quotes them yet.

`optimization.budget` and `optimization.plan` are reserved for an optimizer
that searches plans on a target and for reproducing a searched plan; setting
either is an error today. Further knobs (weight layout, activation LUTs, code
specialization, placement) will join the same section.

## How auto resolves

| Knob | `latency` with `allow_approximate: true` | `latency` | `balanced` | `size` | `accuracy` |
| --- | --- | --- | --- | --- | --- |
| `accumulation` | `fast` | `precise` | `precise` | `precise` | `precise` |
| `kernel` | `specialized` | `specialized` | `specialized` | `generic` | `specialized` |

`auto` never picks an approximate value unless `allow_approximate` is `true`. A value
you set explicitly, for the model or an operator, is used as given: setting
`accumulation: fast` is your consent to its numerics, and the plan records it
as approximate.

## Accumulation

`accumulation` applies to FP16 1x1 convolutions and fully connected layers.
With `precise`, the weights keep the standard layout and the kernels use
ns-cmsis-nn's default accumulation: today FP16 partial sums in vector lanes on
MVE, and blockwise FP16 partials folded into FP32 once the library provides
it. With `fast`, the weights are packed so output channels become vector lanes
and each output accumulates in one long FP16 chain, which is approximate
relative to `precise`. On MVE both accumulate in FP16 today; the difference is
how many partial sums the kernel keeps. Packing pays off only on an MVE
FP16 target with at least 4 rows sharing the weights and at least one full lane
group of 8 outputs; other layers keep the standard layout, and the report marks
the knob not applicable there with the reason. See
[packed float weights](https://ambiqai.github.io/helia-aot/guide/precision/#packed-float-weights) for
the layout and the accuracy analysis. FP32 weights pack where it pays off
regardless of this knob, and keep FP32 accumulation.

Let `auto` choose for latency:

```yaml
optimization:
  goal: latency
  allow_approximate: true
```

Or set the knob for the whole model, and override it per operator type or id:

```yaml
optimization:
  accumulation: fast
operators:
  - type: FULLY_CONNECTED
    id: "12"
    optimization:
      accumulation: precise
```

`accumulation` replaces the `fp16_accuracy` operator attribute of an earlier
development build; that name was never released, so it has no alias.

## Kernel

`kernel` applies to int8 convolutions and depthwise convolutions on MVE
targets; other targets keep today's kernels. With `specialized`, a layer calls
the ns-cmsis-nn 7.37.0 direct entry whose shape rule it meets:

- `arm_convolve_s8_small_cin` for 1 to 3 input channels;
- `arm_convolve_s8_3x3_c16_s1` for 3x3 kernels over 16 channels with stride 1;
- `arm_depthwise_conv_s8_opt_3x3_c64_s1` and `arm_depthwise_conv_s8_opt_3x3` for
  3x3 depthwise layers;
- `arm_depthwise_conv_s8_opt_planar` where the planar rule holds;
- `arm_depthwise_conv_s8_opt_channelwise` for every other layer the optimized
  depthwise kernel takes.

With `generic`, no layer calls a 7.37.0 direct entry: a layer that would take
one calls `arm_convolve_s8` or `arm_depthwise_conv_s8_opt` instead. The
existing shape-specific kernels (1x1, 1xN, and `arm_depthwise_conv_3x3_s8` on
DSP targets) apply with either value. Both values give identical output. Each
layer calls one entry directly, with no run-time dispatch.

`generic` is not always the smaller module. It avoids linking several direct
entries next to the general ones, but `arm_depthwise_conv_s8_opt` carries both
its planar and channel paths, and a 3x3 depthwise layer needs a larger scratch
buffer with it. Code size with `specialized` against `generic`, linked for the
Cortex-M55 with unused sections removed:

- the MLPerf Tiny keyword-spotting model is 1928 B smaller, with 1120 B less
  arena;
- the visual wake words model is 6520 B larger;
- the image classification model is 2652 B larger.

`goal: size` picks `generic`; measure your model if the difference matters. A module that calls
a direct entry requires ns-cmsis-nn 7.37.0 or later; other modules keep their
floor. The report names the entry every operator calls under `entry`.

## Constant placement

Constants are read-only, so their cold copy stays in non-volatile memory, set
with `memory`. `constant_destination_memory` adds a runtime copy in a faster
memory, which the generated `<prefix>_model_init` fills once. Staging weights
into DTCM removes the SRAM weight-read stalls of weight-bound layers such as
fully connected and conv1d; activation-bound models change little. DTCM is
shared with your application's stack and data, so the report states how much
of it the module leaves free, within `memory.constraints`. See
[memory placement](https://ambiqai.github.io/helia-aot/guide/memory-placement/).

```yaml
memory:
  tensors:
    - type: constant
      attributes: {memory: MRAM, constant_destination_memory: DTCM}
```

## The plan and the report

`<prefix>_plan.json` sits next to the module and holds the choices. It carries
a `schema_version`, the heliaAOT version and the ns-cmsis-nn version the module
requires, the model it was made from (file name, sha256 and subgraph), the
target (name, core, capabilities, FP16 support), the goal and approximation
gate, and one entry per operator under `operators`.

Each operator has a stable `id`, `"<subgraph>:<output tensor name>"`, taken
from the name the model gives the operator's first output tensor, so it stays
the same when other operators are added or removed. Its position in execution
order is kept in `index`, and its AIR node id in `node_id`. Repeated names get
`#1`, `#2` in operator order, and an operator without a name falls back to
`"<subgraph>:<op_type>@<position>"`; `id_source` says which applied. A tensor
name can itself contain `:` (for example `0:StatefulPartitionedCall_1:0`), so
a parser splits these ids on their first colon only. An id with
`id_source: node_id` (see below) is the opaque AIR node id and is not split.
`operators[].id` rules match the stable id as well as the node id, so a plan's
ids can be written back as rules. A stable id that would also select another
operator (another operator's node id, or a stable id spelled twice) is not
matched; the plan then uses the node id as `id`, with `id_source: node_id`,
and the conversion warns.

For each knob of an operator the plan records the value used, the value
requested, `resolved_by` (`explicit`, or `auto_table` when `auto` chose it
from the goal; `auto_heuristic` and `plan` are reserved), whether the value is
approximate where it applies, and the `source`: `kind` is `operator` (with
`rule`, the position, type and id of the winning `operators` entry), `model`
(set under `optimization`, even to `auto`), or `default` (`plan` is reserved).
Under `fixed`, it records choices that change numerics or speed but are not
knobs yet, with what set them (`attribute`, `default`, `heuristic` or `knob`;
`plan` is reserved): `RSQRT` `lut_mode` (INT16 only),
`HARD_SWISH` `use_lut` and `compat_variant` (INT8), and the float weight
layout of 1x1 convolutions and fully connected layers (`weight_layout`: set
by `accumulation` for FP16, except where `fast` leaves the standard layout
because packing would not pay off, and by a heuristic otherwise).

`<prefix>_report.json` explains the plan and does not change it: each
operator's MACs, the alternatives of each knob and whether they apply on this
target (and why not), the measurements for this target and core keyed by
setting, the constant placement facts, and the hints the results print. Its
content may change between releases; the choices are in the plan.

Both files carry a `schema_version`. Adding an optional field, a knob, a knob
value or an enum member keeps it; renaming or removing a field, or changing
what one means, increments it. The JSON schemas are exported with the
documentation data (`plan-schema.json`, `report-schema.json`) and list every
knob with its values, so they belong to the heliaAOT release that wrote a file
(`generator.helia_aot`): validate a file against that release's schema rather
than pinning one. `helia_aot.optimization` reads both files back as typed
models. Some values and fields are reserved and empty today: `plan` as a
source, `optimization.knobs` for model-level choices and `tensors` for
per-tensor choices.

Knobs are set only under `optimization`, for the model or in an `operators`
rule; a knob name under `attributes` is rejected. `optimization.budget` and
`optimization.plan` are YAML-only while they are reserved.

A hint appears only when its alternative applies and a single setting would
select it. A knob hint lists the operators where the knob applies that do not
run the value `auto` picks for `goal: latency` with `allow_approximate: true`
(for accumulation, operators running `precise` that `fast` would pack),
leaving out operators an `operators` rule pins. The placement hint needs constants outside TCM that all come from one
cold memory and fit in the free DTCM. Share of MACs counts convolution,
depthwise and fully connected layers; it is not a share of run time unless you
profile the model.
