Skip to content
heliaAOT
HELIA HUB

Precision

heliaAOT compiles the tensor types and quantization parameters already present in the LiteRT file. Training, calibration and precision changes happen upstream. This does not promise bit-identical interpreter output: validate the generated module with known inputs and a predetermined numerical acceptance criterion.

Check three things before choosing a path: the model tensor types, operator coverage and the target/library build requirements. Precision affects memory use and kernel selection, but the complete model must be measured to compare size or speed.

For quantized compute paths, A8W8 means 8-bit activations and 8-bit weights; A16W8 means 16-bit activations and 8-bit weights. These labels describe a combination of tensor roles. The int8 and int16 types below identify individual tensors; biases, accumulators and state can use other types.

Path What the model carries Where it runs Reach for it when
int8 Quantized data tensors; weights and metadata have their own constraints Every registered target, where each operator accepts the model Start with the supplied KWS example or an already validated quantized model
int16 16-bit activations, for the operators that declare an int16 kernel Every registered target, per operator int8 calibration will not hold the dynamic range
FP16 Half-precision data tensors; index, mask and other metadata retain their required types Native compute needs supports_fp16; storage/copy paths can differ The model needs half-precision computation and each operator/target supports it
FP32 float32 data tensors, with separate index, mask and state requirements Targets with a hardware floating-point unit You need reference numerics, or you are bringing a model up before quantizing it

Wider tensors increase storage per element. Total flash, arena use and cycles also depend on the model, selected kernels and target; precision alone does not predict their ordering. Check each operator’s tensor and target requirements.

The first-deployment example uses int8, and every registered platform has integer kernel paths. Check the operator catalog for the particular graph. An int8 element takes fewer bytes than an int16 or float32 element, but total constants, scratch and code also depend on shapes, operator choices and workspace requirements. Compare the generated memory reports rather than predicting whole-module size from precision alone.

The compiler reads quantization parameters from the flatbuffer and embeds them in generated code. Kernels can still rescale intermediate values during inference. Calibration happens upstream; check the exported scales and zero points when investigating a numerical mismatch.

Two int8 details show up in the catalog’s Restrictions column often enough to be worth knowing before you export:

  • Several operators require a specific zero point. LOGISTIC wants an int8 output at zero-point -128, for example, because the kernel assumes it.
  • Weights on the recurrent path have to be symmetric, which means zero-point 0.

Check the exported tensor parameters rather than assuming the exporter selected the required form. Conversion validates operator-specific restrictions.

int16 where the operator has an int16 kernel

Section titled “int16 where the operator has an int16 kernel”

int16 widens the activations while keeping integer arithmetic, which is the usual answer when a model has a layer whose range int8 cannot hold. It is not a whole-model switch that heliaAOT applies: it is a property of the tensors in the model you hand it. Check whether each operator accepts int16 data tensors; that may select a kernel, inline code or a copy route.

The Data types column of the operator catalog is that list. Check every operator your model uses before you export an int16 version of it, and read the Restrictions column in the same row: int16 brings its own conditions, such as operators that accept int16 only with a zero point of 0, and operators whose output type is constrained by the input type.

The catalog summarizes data types; index, mask and recurrent state tensors have separate roles. In particular, an int16 LSTM cell state does not establish an int16-input/output LSTM path.

Operators validate types and other conditions per node, using declarations or per-instance checks. A dtype match alone does not establish support for its shape, quantization, attributes or target. Mixed precision also requires valid conversion boundaries: DEQUANTIZE accepts int8, int16 or float16 input and produces float32; QUANTIZE accepts float32 to int8/int16 or requantizes within the same integer width. These catalog types are not interchangeable input/output pairs.

A declared float API or C type dependency requires the corresponding library switch below. Operators using native FP16 kernels also check the platform’s FP16 capability. Byte copies and storage conversions have different needs: EXPAND_DIMS copies bytes, while FP16-to-FP32 DEQUANTIZE can widen stored values in software on a target without native FP16 kernels.

How heliaAOT works describes the Resolve stage these checks belong to, and Troubleshooting lists the error classes they raise.

heliaAOT can compile a model whose tensors are float16 or float32 and emit float calls to heliaCORE, the Ambiq fork of CMSIS-NN, instead of quantized paths where supported. Some operations use inline arithmetic, lookup loops, byte copies or alias no-ops instead. A library call can itself select a vector, DSP or scalar implementation depending on the shape, target and build flags; float support is not a count of optimized kernels.

Format Required core feature Where to check
FP32 Hardware single-precision FPU Targets
FP16 Armv8.1-M MVE-FP, half-precision vector floating point The FP16 capability in the Targets reference

Emitting _f16 kernels for a core without MVE-FP would either fail to build or misbehave silently, so heliaAOT refuses at conversion time. Operators that call the shared native-FP16 support check reject a float16 tensor when the platform’s supports_fp16 check fails, naming the platform and core. This gate does not apply universally to FP16 storage or byte movement. The platform’s supports_fp16 property accepts a compatible explicit FP16 capability and also infers support from a normalized cortex-m55 CPU when that capability is absent. Omitting that capability from a custom Cortex-M55 definition therefore does not disable every native half-precision path. Extra capability fields on registered target names are ignored; use the custom-target mechanism when changing hardware declarations.

FP32 has no such gate: any target with a floating-point unit can run it.

LiteRT’s built-in FP16 conversion is a storage optimization. It keeps weight constants as FP16 but inserts a dequantize in front of each layer and leaves activations and the model’s inputs and outputs in FP32.

A model produced that way does not exercise the FP16 kernels. heliaAOT sees the dequantize followed by an FP32 kernel, which is exactly what it emits. For weightless operators there is nothing to convert at all, so they arrive as plain FP32.

To exercise native FP16 computation, the relevant operator’s data inputs, weights, biases and outputs must meet its FP16 requirements. Metadata/index tensors can retain other types, and a mixed graph can have valid conversion boundaries. Inspect the actual tensor roles and emitted path rather than classifying the entire model from its filename.

Float 1x1 convolutions (arm_convolve_1x1_f16/f32) and float fully connected layers can store their weights in ns-cmsis-nn’s NT_N_PACKED layout, [ceil(C_OUT / L)][C_IN][L] with L = 8 for FP16 and 4 for FP32, so each output channel is a vector lane. On Apollo510 a pointwise layer (width 256, 32 channels) measured 3.22x faster at FP16 and 2.19x at FP32 packed.

  • When it applies. FP32 packs by default; FP16 packs only when the accumulation optimization knob resolves to fast (see Performance and accuracy options). Either way the weights are packed only on an MVE target, with at least 4 rows sharing them (output positions, or batch rows for fully connected) and at least one full lane group of outputs. A batch-1 fully connected layer keeps the standard layout, which runs fewer instructions there. The library floor is unchanged (7.35.0).
  • Accuracy. FP16 accumulates in FP16 lanes on both layouts. From 32 input channels up, the standard path keeps several partial sums while the packed path keeps one per output. A simulation puts the packed error at roughly 4-5x the standard path’s, about 0.018 at 512-1024 input channels for outputs near 1, which is why auto resolves to precise unless you allow approximate values. FP32 keeps FP32 accumulation on both layouts.
  • Storage. Output channels are padded to a multiple of L, so a packed layer stores up to 7 (FP16) or 3 (FP32) padding rows of weights.
  • Tensor ids and placement. Weights that only one operator reads are packed in place and keep their tensor id. Weights shared with other operators get one packed copy, <id>_packed (or <id>_packed_N when that id is taken); the original stays for any reader that does not pack and is dropped when none is left. memory.tensors rules name the original id and also place the copy; a rule on the _packed id itself is not applied. Weights that are a model input or a variable are written at run time and keep the standard layout.

Start with the generated module’s dependency requirements and enable the needed precisions on the kernel library before adding the generated module. A compiler flag on application code alone cannot add missing library objects.

The float kernel objects exist only if ns-cmsis-nn itself was built with them, so the switch belongs on the library, not on the generated module.

Build system Switch
CMake and NSX ARM_NN_ENABLE_F32, ARM_NN_ENABLE_F16, set before adding the ns-cmsis-nn subdirectory
Zephyr CONFIG_NS_CMSIS_NN_ENABLE_F32, CONFIG_NS_CMSIS_NN_ENABLE_F16 in the application’s prj.conf

On the standalone CMake and NSX backends the module does not define these itself. It inherits them from the kernel library’s usage interface and asks the library what it was built with, through a capability query that ns-cmsis-nn 7.35.0 and later provide. A linked kernel target without that query fails configure with a hint to load the library’s normal CMake entry point. The library’s answer overrides a conflicting cache variable.

Only the precisions the operators actually declare are required. An FP16-only module can have F16 on and F32 off; an FP32-only module does not need F16. A module with no float kernel or float type dependency at all - an FP32 reshape, a metadata-only shape query - needs neither. Half-precision utilities that bind the half-precision type still require F16 even when they only move bytes.

On standalone CMake a parent may bring its own kernel target instead. The module then links nothing and defines only the switches its operators require. Set a required switch on, or leave it unset to take that default; setting a required switch off fails configure.

Coverage is per operator, per element type and per target, not per model. The operator catalog is the table: every operator heliaAOT lowers, its dtype summary, declared float API or type dependency, and the oldest library that float path builds against. The module-wide kernel library floor is on the Versions page.

The conversion itself checks whether each remaining node has an applicable lowering. Then build and validate the generated module with the intended library and target. Successful conversion alone does not establish optimized execution or numerical correctness. Use the output verification workflow with the intended library, compiler and execution environment. Keep the acceptance threshold fixed while investigating a mismatch; increasing it to fit observed errors is not validation.