# Precision

heliaAOT compiles the tensor types and quantization parameters already present
in the LiteRT file. Training, calibration and precision changes happen
upstream. This does not promise bit-identical interpreter output: validate the
generated module with known inputs and a predetermined numerical acceptance
criterion.

Check three things before choosing a path: the model tensor types, operator
coverage and the target/library build requirements. Precision affects memory
use and kernel selection, but the complete model must be measured to compare
size or speed.

For quantized compute paths, **A8W8** means 8-bit activations and 8-bit weights;
**A16W8** means 16-bit activations and 8-bit weights. These labels describe a
combination of tensor roles. The `int8` and `int16` types below identify
individual tensors; biases, accumulators and state can use other types.

## The four paths

| Path | What the model carries | Where it runs | Reach for it when |
| --- | --- | --- | --- |
| int8 | Quantized data tensors; weights and metadata have their own constraints | Every registered target, where each operator accepts the model | Start with the supplied KWS example or an already validated quantized model |
| int16 | 16-bit activations, for the operators that declare an int16 kernel | Every registered target, per operator | int8 calibration will not hold the dynamic range |
| FP16 | Half-precision data tensors; index, mask and other metadata retain their required types | Native compute needs `supports_fp16`; storage/copy paths can differ | The model needs half-precision computation and each operator/target supports it |
| FP32 | `float32` data tensors, with separate index, mask and state requirements | Targets with a hardware floating-point unit | You need reference numerics, or you are bringing a model up before quantizing it |

Wider tensors increase storage per element. Total flash, arena use and cycles
also depend on the model, selected kernels and target; precision alone does not
predict their ordering. Check each operator's tensor and target requirements.

## int8 is the default path

The first-deployment example uses int8, and every registered platform has
integer kernel paths. Check the [operator catalog](https://ambiqai.github.io/helia-aot/reference/operators/)
for the particular graph. An int8 element takes fewer bytes than an int16 or
float32 element, but total constants, scratch and code also depend on shapes,
operator choices and workspace requirements. Compare the generated memory
reports rather than predicting whole-module size from precision alone.

The compiler reads quantization parameters from the flatbuffer and embeds them
in generated code. Kernels can still rescale intermediate values during
inference. Calibration happens upstream; check the exported scales and zero
points when investigating a numerical mismatch.

Two int8 details show up in the catalog's Restrictions column often enough to
be worth knowing before you export:

- Several operators require a specific zero point. `LOGISTIC` wants an int8
  output at zero-point -128, for example, because the kernel assumes it.
- Weights on the recurrent path have to be symmetric, which means zero-point 0.

Check the exported tensor parameters rather than assuming the exporter selected
the required form. Conversion validates operator-specific restrictions.

## int16 where the operator has an int16 kernel

int16 widens the activations while keeping integer arithmetic, which is the
usual answer when a model has a layer whose range int8 cannot hold. It is not a
whole-model switch that heliaAOT applies: it is a property of the tensors in
the model you hand it. Check whether each operator accepts int16 data tensors;
that may select a kernel, inline code or a copy route.

The Data types column of the [operator catalog](https://ambiqai.github.io/helia-aot/reference/operators/)
is that list. Check every operator your model uses before you export an int16
version of it, and read the Restrictions column in the same row: int16 brings
its own conditions, such as operators that accept int16 only with a zero point
of 0, and operators whose output type is constrained by the input type.

The catalog summarizes data types; index, mask and recurrent state tensors
have separate roles. In particular, an int16 LSTM cell state does not establish
an int16-input/output LSTM path.

## Where the decision is enforced

Operators validate types and other conditions per node, using declarations or
per-instance checks. A dtype match alone does not establish support for its
shape, quantization, attributes or target. Mixed precision also requires valid
conversion boundaries: `DEQUANTIZE` accepts int8, int16 or float16 input and
produces float32; `QUANTIZE` accepts float32 to int8/int16 or requantizes within
the same integer width. These catalog types are not interchangeable input/output
pairs.

A declared float API or C type dependency requires the corresponding library
switch below. Operators using native FP16 kernels also check the platform's
FP16 capability. Byte copies and storage conversions have different needs:
`EXPAND_DIMS` copies bytes, while FP16-to-FP32 `DEQUANTIZE` can widen stored
values in software on a target without native FP16 kernels.

[How heliaAOT works](https://ambiqai.github.io/helia-aot/guide/how-it-works/) describes the Resolve
stage these checks belong to, and
[Troubleshooting](https://ambiqai.github.io/helia-aot/guide/troubleshooting/) lists the error classes
they raise.

## Float support

:::caution
Floating-point operator support is experimental. The surface, the set of
supported operators and the selection behaviour may change without notice.
Prefer quantized int8 or int16 for anything shipping.
:::

heliaAOT can compile a model whose tensors are `float16` or `float32` and emit
float calls to heliaCORE, the Ambiq fork of CMSIS-NN, instead of quantized
paths where supported. Some operations use inline arithmetic, lookup loops,
byte copies or alias no-ops instead. A library call can itself select a vector,
DSP or scalar implementation depending on the shape, target and build flags;
float support is not a count of optimized kernels.

### What the hardware has to provide

| Format | Required core feature | Where to check |
| --- | --- | --- |
| FP32 | Hardware single-precision FPU | [Targets](https://ambiqai.github.io/helia-aot/guide/targets/) |
| FP16 | Armv8.1-M MVE-FP, half-precision vector floating point | The `FP16` capability in the [Targets reference](https://ambiqai.github.io/helia-aot/reference/targets/) |

### The platform gate

Emitting `_f16` kernels for a core without MVE-FP would either fail to build or
misbehave silently, so heliaAOT refuses at conversion time. Operators that call the shared native-FP16 support check reject a `float16`
tensor when the platform's `supports_fp16` check fails, naming the platform and core.
This gate does not apply universally to FP16 storage or byte movement. The
platform's `supports_fp16` property accepts a compatible explicit `FP16`
capability and also infers support from a normalized `cortex-m55` CPU when that
capability is absent. Omitting that capability from a custom Cortex-M55
definition therefore does not disable every native half-precision path. Extra
capability fields on registered target names are ignored; use the
[custom-target mechanism](https://ambiqai.github.io/helia-aot/guide/targets/#overriding-a-platform) when
changing hardware declarations.

FP32 has no such gate: any target with a floating-point unit can run it.

### What is and is not a float model

LiteRT's built-in FP16 conversion is a storage optimization. It keeps weight
constants as FP16 but inserts a dequantize in front of each layer and leaves
activations and the model's inputs and outputs in FP32.

A model produced that way does not exercise the FP16 kernels. heliaAOT sees the
dequantize followed by an FP32 kernel, which is exactly what it emits. For
weightless operators there is nothing to convert at all, so they arrive as
plain FP32.

To exercise native FP16 computation, the relevant operator's data inputs,
weights, biases and outputs must meet its FP16 requirements. Metadata/index
tensors can retain other types, and a mixed graph can have valid conversion
boundaries. Inspect the actual tensor roles and emitted path rather than
classifying the entire model from its filename.

:::note
A CMSIS-NN kernel is typed by a single dtype suffix. There is no kernel that
mixes FP16 weights with FP32 activations; that combination exists only as the
LiteRT storage trick above and has to be resolved before a kernel runs.
:::

### Packed float weights

Float 1x1 convolutions (`arm_convolve_1x1_f16/f32`) and float fully connected
layers can store their weights in ns-cmsis-nn's `NT_N_PACKED` layout,
`[ceil(C_OUT / L)][C_IN][L]` with L = 8 for FP16 and 4 for FP32, so each output
channel is a vector lane. On Apollo510 a pointwise layer (width 256, 32
channels) measured 3.22x faster at FP16 and 2.19x at FP32 packed.

- **When it applies.** FP32 packs by default; FP16 packs only when the
  `accumulation` optimization knob resolves to `fast` (see
  [Performance and accuracy options](https://ambiqai.github.io/helia-aot/guide/options/)). Either way
  the weights are packed only on an
  MVE target, with at least 4 rows sharing them (output positions, or batch
  rows for fully connected) and at least one full lane group of outputs. A
  batch-1 fully connected layer keeps the standard layout, which runs fewer
  instructions there. The library floor is unchanged (7.35.0).
- **Accuracy.** FP16 accumulates in FP16 lanes on both layouts. From 32 input
  channels up, the standard path keeps several partial sums while the packed
  path keeps one per output. A simulation puts the packed error at roughly 4-5x
  the standard path's, about 0.018 at 512-1024 input channels for outputs near
  1, which is why `auto` resolves to `precise` unless you allow approximate
  values. FP32 keeps FP32 accumulation on
  both layouts.
- **Storage.** Output channels are padded to a multiple of L, so a packed layer
  stores up to 7 (FP16) or 3 (FP32) padding rows of weights.
- **Tensor ids and placement.** Weights that only one operator reads are packed
  in place and keep their tensor id. Weights shared with other operators get
  one packed copy, `<id>_packed` (or `<id>_packed_N` when that id is taken); the
  original stays for any reader that does not pack and is dropped when none is
  left. `memory.tensors` rules name the original id and also place the copy; a
  rule on the `_packed` id itself is not applied. Weights that are a model input
  or a variable are written at run time and keep the standard layout.

### The library build switch

Start with the generated module's dependency requirements and enable the needed
precisions on the kernel library before adding the generated module. A compiler
flag on application code alone cannot add missing library objects.

The float kernel objects exist only if ns-cmsis-nn itself was built with them,
so the switch belongs on the library, not on the generated module.

| Build system | Switch |
| --- | --- |
| CMake and NSX | `ARM_NN_ENABLE_F32`, `ARM_NN_ENABLE_F16`, set before adding the ns-cmsis-nn subdirectory |
| Zephyr | `CONFIG_NS_CMSIS_NN_ENABLE_F32`, `CONFIG_NS_CMSIS_NN_ENABLE_F16` in the application's `prj.conf` |

On the standalone CMake and NSX backends the module does not define these itself. It inherits them from the kernel library's usage interface and asks the library what it was built with, through a capability query that ns-cmsis-nn 7.35.0 and later provide.
A linked kernel target without that query fails
configure with a hint to load the library's normal CMake entry point. The
library's answer overrides a conflicting cache variable.

Only the precisions the operators actually declare are required. An FP16-only
module can have F16 on and F32 off; an FP32-only module does not need F16. A
module with no float kernel or float type dependency at all - an FP32 reshape,
a metadata-only shape query - needs neither. Half-precision utilities that bind
the half-precision type still require F16 even when they only move bytes.

On standalone CMake a parent may bring its own kernel target instead. The
module then links nothing and defines only the switches its operators require.
Set a required switch on, or leave it unset to take that default; setting a
required switch off fails configure.

:::caution
Zephyr checks F32 for every module with a float dependency, including
FP16-only modules, and checks F16 as well when required. The neuralSPOT
integration instructions also ask for F32 on FP16-only modules. Those two
backends have not adopted independent precision selection.

ns-cmsis-nn 7.35.0 and later also reject the removed NSX float-enable switch
names at configure time. Use `ARM_NN_ENABLE_F32` and `ARM_NN_ENABLE_F16`; the
Zephyr symbols are unchanged.
:::

## Checking a model before you convert

Coverage is per operator, per element type and per target, not per model. The
[operator catalog](https://ambiqai.github.io/helia-aot/reference/operators/) is the table: every
operator heliaAOT lowers, its dtype summary, declared float API or type
dependency, and the oldest library that float path builds against. The module-wide kernel library floor is on the
[Versions page](https://ambiqai.github.io/helia-aot/reference/versions/).

The conversion itself checks whether each remaining node has an applicable
lowering. Then build and validate the generated module with the intended
library and target. Successful conversion alone does not establish optimized
execution or numerical correctness. Use the
[output verification workflow](https://ambiqai.github.io/helia-aot/getting-started/validate/) with the
intended library, compiler and execution environment. Keep the acceptance
threshold fixed while investigating a mismatch; increasing it to fit observed
errors is not validation.
