Migrating from heliaRT
If your application already uses heliaRT, the main migration is from an
interpreter-managed model to generated source and a fixed memory plan. Begin
with the same .tflite model, then check that its operators, export form and
tensor roles are supported by heliaAOT.
What stays the same
Section titled “What stays the same”Training, preprocessing and quantization remain upstream responsibilities. heliaAOT reads tensor types, scales and zero points from the flatbuffer; it does not retrain or recalibrate it. A compatible existing model can be converted without re-exporting. An unsupported construct may require an export change.
Both products can use Ambiq kernels, but that does not establish bit-identical outputs. heliaAOT can also emit arithmetic loops, copies and aliases. Validate the generated module against known outputs before comparing performance.
What changes
Section titled “What changes”| Concern | Migration to heliaAOT |
|---|---|
| Firmware artifact | Compile generated sources and constants instead of supplying a flatbuffer to an interpreter. |
| Operator selection | Conversion resolves each node; there is no firmware operator resolver to populate. |
| Inference API | Initialize a generated context, populate its input buffers, run, then read output buffers. |
| Memory | Use generated buffers by default, or bind caller-owned regions before initialization. |
| Verification | Run the generated golden-data test in the intended firmware environment. |
The artifact
Section titled “The artifact”The module contains public headers in includes-api/, sources in src/, build
integration files and a README describing model I/O and memory. Keep the model,
configuration and compiler version as the source of regeneration. Do not carry
hand edits into generated files.
Generated files explains the artifact and provides an input-copy/run/output-copy wrapper for the walkthrough model.
The C API
Section titled “The C API”#include "aot_model.h"
static aot_model_context_t ctx = {0};
int32_t start_model(void){ return aot_model_init(&ctx);}Check start_model() before accessing ctx.inputs. Copy the prepared tensor
bytes into ctx.inputs[i].data, call aot_model_run(&ctx), check its status and
then consume ctx.outputs[i].data. The descriptor’s size is bytes, and its
scale and zero_point describe quantization. For quantized values,
real_value = (q - zero_point) * scale.
Use the model’s input/output order and dtype, not assumptions from the pointer’s storage type. Preserve your existing preprocessing contract and compare one known input before connecting live data.
Who owns the arena
Section titled “Who owns the arena”With default memory.allocate_arenas: true, the module declares its buffers.
With it false, bind every region using aot_bind_arena or aot_bind_arenas
before aot_model_init. Sizes, alignment and constant sidecars are part of that
contract; Memory covers them.
Calls must be serialized: arena bindings and hydration state are module-global, so creating two context structs does not create two independent concurrent instances. Sharing storage between models also requires preserving any state that must survive the switch.
Internal recurrent state persists between aot_model_run calls and resets on
aot_model_init. A model with explicit state outputs/inputs needs caller-managed
carry. See Stateful models.
Build integration
Section titled “Build integration”module.type |
Integration entry point |
|---|---|
neuralspot |
module.mk |
zephyr |
zephyr/module.yml, Kconfig and CMakeLists.txt |
cmake |
CMakeLists.txt |
nsx |
CMakeLists.txt and nsx-module.yaml |
cmsis_pack |
.pdsc manifest, optionally packaged as .pack |
Select the format matching the firmware project and regenerate. Recheck kernel library version, compile options, logging and linker placement using Firmware integration. Format support does not mean every existing heliaRT project can accept the module without changes.
Testing
Section titled “Testing”Set test.enabled: true and provide an NPZ containing input_N and output_N
arrays from a known model execution. The generated test embeds these bytes and
compares final output with the configured tolerance. Check init/run statuses and
per-output comparison records with verification enabled; a success message with
verification disabled proves no output match.
On-device validation gives the complete workflow. Evaluate representative data separately from this smoke test.
Before you move a model
Section titled “Before you move a model”- Confirm it is a LiteRT flatbuffer with shapes resolvable at conversion time.
- Check operator, tensor-role and target constraints in the catalog, then run conversion.
- Identify persistent or explicit state and define reset boundaries.
- Build with the intended library, toolchain and memory mapping.
- Compare known outputs before connecting live input or profiling.
A first pass
Section titled “A first pass”Use the Getting started path to establish the conversion/build/test workflow with the supplied KWS model. Then substitute the model you ship and its fixtures. Once both engines meet the numerical contract, use the measurement guide to compare recorded runs under controlled conditions.