How heliaAOT works
heliaAOT turns an exported LiteRT model into a model-specific C module. Your development machine does the graph processing, kernel selection and memory planning. Your firmware initializes the generated module, supplies input data and calls it to run inference.
You still build and link firmware for your board. The generated module uses kernel-library code and a small amount of generated execution support; it does not need a LiteRT interpreter to parse and plan the model on the device.
From model and settings to a module
Section titled “From model and settings to a module”- Your model and settings01 · Input
- heliaAOT02 · Compile on your development machine
- Model code and artifacts03 · Output
- Operator codeGenerated C and selected kernel calls
- Memory arenasConstants, persistent state and scratch
- Execution and APIsInitialization and a generated model schedule
- PackagingEmbedded constants or configured sidecar blobs
- DocumentationModule README and optional HTML or memory reports
- Build integrationCMake, Zephyr, neuralSPOT, NSX or CMSIS-Pack
This is a conceptual build-time view. The output is a module ready for firmware integration, not a board executable: your application still compiles and links it with the selected kernels and platform support. Useful defaults mean you do not need to specify every deployment setting.
Training and model export happen upstream. The compiler accepts LiteRT flatbuffers (.tflite or .litert); it does not train a network or automatically convert arbitrary models between integer and float precision. Start with the first-conversion walkthrough for a supplied model and configuration.
Before deployment
Section titled “Before deployment”The conversion makes decisions that would otherwise need runtime machinery or hand-written integration work.
These blocks expand the heliaAOT box above into user-relevant decisions, rather than specifying an exact compiler pass order. The compiler architecture gives that detailed sequence. The controls table below links each decision to optional settings.
Analyze and transform the model
Section titled “Analyze and transform the model”The compiler reads the selected subgraph, applies registered model hooks and runs graph transforms. The built-in transform pipeline folds static shape expressions, rewrites eligible depthwise convolutions, prunes identity operations and rewrites eligible transpose convolutions. These passes are enabled by default; each only changes patterns it supports.
Shapes must resolve before code generation. A successful conversion means the compiler could lower this model for the selected settings; it does not prove the firmware produces acceptable outputs on your device.
Select an implementation for each operation
Section titled “Select an implementation for each operation”Operators check their supported dtypes, layouts, shapes, quantization and platform requirements. Lowering may call a heliaCORE/ns-cmsis-nn kernel, emit inline code, copy data or eliminate a no-op. The operator catalog describes compatibility; an entry is not a promise that every shape and dtype uses an optimized kernel.
Where an operator offers several kernels, applicability predicates and declared priorities select one. Selection is not an on-device search or a measured cost model. Precision explains the additional float and platform restrictions.
Lay out memory before execution
Section titled “Lay out memory before execution”The compiler removes unused tensors, interns eligible identical constants and assigns tensors to arenas. It reuses scratch slots when tensor lifetimes do not overlap. Persistent storage carries state across invocations; constant storage supplies weights and other fixed values. The default scratch planner is greedy.
Memory planning gives the module a known layout and arena requirements. Your firmware still has to provide the corresponding memory and linker placement, with space for the rest of the application. Read the memory reports before choosing placement or another planner.
Write a module you can build
Section titled “Write a module you can build”The result includes model/context/tensor code, operator implementation files, constants, headers and the build integration selected by module.type. The supported formats are neuralSPOT, Zephyr, CMake, NSX and CMSIS-Pack. Optional generated tests and reports help inspect and validate the result.
Changing build format does not replace firmware integration or validation. Keep the generated module together with its model, configuration and dependency versions. The generated-files walkthrough shows what to keep and where to start.
On the device
Section titled “On the device”- InitializePrepare storage, constants and operators
- Supply inputs and runUse the model input format and check status
- Consume outputsRead results before storage is reused
Use the generated model header for your module’s exact names. With the default module.prefix: aot, the public entry points are aot_model_init() and aot_model_run().
- Prepare storage and initialize. Start with a zero-initialized context. In application-owned arena mode, bind every required arena first. Call
aot_model_init()and stop on a nonzero status. Initialization resolves tensor pointers, resets persistent state, hydrates staged constants when configured, and initializes operators. - Write inputs and invoke. Use the initialized context’s input views and the model’s dtype, shape and quantization. Fill input bytes before calling
aot_model_run(). Check its status before trusting outputs. - Consume outputs, then repeat. Read or copy output data before the next invocation or reinitialization can reuse its storage. Keep the initialized module for successive samples when state should carry. Reinitialize when you deliberately need the model’s reset state.
Follow firmware integration for a complete application path and output validation before using inference results.
What remains at runtime
Section titled “What remains at runtime”The firmware still contains the selected kernels, constants, tensor descriptors and pointer setup, memory arenas, status handling and the generated model’s execution schedule. Staged constants add a copy into their runtime arena during initialization.
The default module.schedule: table runs a generated operator table and supports per-node callbacks. static emits direct calls and omits that callback seam. Both use the same public model API. Neither schedule requires a general model parser on the device; compare the resulting firmware before claiming a size or latency improvement.
Start with defaults, add control for a reason
Section titled “Start with defaults, add control for a reason”A first conversion already applies default transforms, selects eligible implementations and plans memory. You do not need to configure every pass or tensor to get a module. Enable generated tests when you need an output-verification harness; tests are not enabled by default.
| When you need to… | Start here | Control to explore |
|---|---|---|
| Reproduce a conversion | Save the model, YAML, package versions and effective settings | Configuration precedence |
| Change a graph rewrite or operator implementation option | Identify the affected operation and validate the changed outputs | Transforms and operator rules |
| Fit a memory budget or choose a bank | Inspect the arena and tensor residency report | Constraints and tensor placement |
| Move constants into writable memory | Account for the source blob, RAM destination and initialization copy | Staged constants |
| Reduce scratch storage | Compare layouts for the same model and constraints | Scratch planners |
| Add instrumentation or compare dispatch shapes | Keep callbacks with table; measure a static alternative separately |
Execution settings |
| Support a custom operation | Supply parsing, validation and code generation, with a trusted oracle | Custom operators |
Transform configuration enables, disables and supplies options to registered passes. It does not reorder them: execution follows registry order. Memory and operator rules likewise change only the matching supported settings; they do not guarantee that every model will fit or become faster.
Compiler details
Section titled “Compiler details”If you are extending the compiler, continue with Compiler architecture for AIR, the six conversion stages, handler order, operator lifecycle and kernel selection. The generated module reference covers the C integration surface and its distinction from implementation files.