Skip to content
heliaAOT
HELIA HUB

Migrating from heliaRT

If your application already uses heliaRT, the main migration is from an interpreter-managed model to generated source and a fixed memory plan. Begin with the same .tflite model, then check that its operators, export form and tensor roles are supported by heliaAOT.

Training, preprocessing and quantization remain upstream responsibilities. heliaAOT reads tensor types, scales and zero points from the flatbuffer; it does not retrain or recalibrate it. A compatible existing model can be converted without re-exporting. An unsupported construct may require an export change.

Both products can use Ambiq kernels, but that does not establish bit-identical outputs. heliaAOT can also emit arithmetic loops, copies and aliases. Validate the generated module against known outputs before comparing performance.

Concern Migration to heliaAOT
Firmware artifact Compile generated sources and constants instead of supplying a flatbuffer to an interpreter.
Operator selection Conversion resolves each node; there is no firmware operator resolver to populate.
Inference API Initialize a generated context, populate its input buffers, run, then read output buffers.
Memory Use generated buffers by default, or bind caller-owned regions before initialization.
Verification Run the generated golden-data test in the intended firmware environment.

The module contains public headers in includes-api/, sources in src/, build integration files and a README describing model I/O and memory. Keep the model, configuration and compiler version as the source of regeneration. Do not carry hand edits into generated files.

Generated files explains the artifact and provides an input-copy/run/output-copy wrapper for the walkthrough model.

#include "aot_model.h"
static aot_model_context_t ctx = {0};
int32_t start_model(void)
{
return aot_model_init(&ctx);
}

Check start_model() before accessing ctx.inputs. Copy the prepared tensor bytes into ctx.inputs[i].data, call aot_model_run(&ctx), check its status and then consume ctx.outputs[i].data. The descriptor’s size is bytes, and its scale and zero_point describe quantization. For quantized values, real_value = (q - zero_point) * scale.

Use the model’s input/output order and dtype, not assumptions from the pointer’s storage type. Preserve your existing preprocessing contract and compare one known input before connecting live data.

With default memory.allocate_arenas: true, the module declares its buffers. With it false, bind every region using aot_bind_arena or aot_bind_arenas before aot_model_init. Sizes, alignment and constant sidecars are part of that contract; Memory covers them.

Calls must be serialized: arena bindings and hydration state are module-global, so creating two context structs does not create two independent concurrent instances. Sharing storage between models also requires preserving any state that must survive the switch.

Internal recurrent state persists between aot_model_run calls and resets on aot_model_init. A model with explicit state outputs/inputs needs caller-managed carry. See Stateful models.

module.type Integration entry point
neuralspot module.mk
zephyr zephyr/module.yml, Kconfig and CMakeLists.txt
cmake CMakeLists.txt
nsx CMakeLists.txt and nsx-module.yaml
cmsis_pack .pdsc manifest, optionally packaged as .pack

Select the format matching the firmware project and regenerate. Recheck kernel library version, compile options, logging and linker placement using Firmware integration. Format support does not mean every existing heliaRT project can accept the module without changes.

Set test.enabled: true and provide an NPZ containing input_N and output_N arrays from a known model execution. The generated test embeds these bytes and compares final output with the configured tolerance. Check init/run statuses and per-output comparison records with verification enabled; a success message with verification disabled proves no output match.

On-device validation gives the complete workflow. Evaluate representative data separately from this smoke test.

  1. Confirm it is a LiteRT flatbuffer with shapes resolvable at conversion time.
  2. Check operator, tensor-role and target constraints in the catalog, then run conversion.
  3. Identify persistent or explicit state and define reset boundaries.
  4. Build with the intended library, toolchain and memory mapping.
  5. Compare known outputs before connecting live input or profiling.

Use the Getting started path to establish the conversion/build/test workflow with the supplied KWS model. Then substitute the model you ship and its fixtures. Once both engines meet the numerical contract, use the measurement guide to compare recorded runs under controlled conditions.