Skip to content
heliaAOT
HELIA HUB

How heliaAOT works

heliaAOT turns an exported LiteRT model into a model-specific C module. Your development machine does the graph processing, kernel selection and memory planning. Your firmware initializes the generated module, supplies input data and calls it to run inference.

You still build and link firmware for your board. The generated module uses kernel-library code and a small amount of generated execution support; it does not need a LiteRT interpreter to parse and plan the model on the device.

  1. Your model and settings01 · Input
  2. heliaAOT02 · Compile on your development machine
  3. Model code and artifacts03 · Output

This is a conceptual build-time view. The output is a module ready for firmware integration, not a board executable: your application still compiles and links it with the selected kernels and platform support. Useful defaults mean you do not need to specify every deployment setting.

Training and model export happen upstream. The compiler accepts LiteRT flatbuffers (.tflite or .litert); it does not train a network or automatically convert arbitrary models between integer and float precision. Start with the first-conversion walkthrough for a supplied model and configuration.

The conversion makes decisions that would otherwise need runtime machinery or hand-written integration work.

  1. Understand the modelOperators, shapes, precision and support
  2. Optimize and select kernelsGraph rewrites and eligible target implementations
  3. Plan memoryConstants, persistent state, scratch and placement
  4. Generate deploymentModel code, integration and enabled reports or tests

These blocks expand the heliaAOT box above into user-relevant decisions, rather than specifying an exact compiler pass order. The compiler architecture gives that detailed sequence. The controls table below links each decision to optional settings.

The compiler reads the selected subgraph, applies registered model hooks and runs graph transforms. The built-in transform pipeline folds static shape expressions, rewrites eligible depthwise convolutions, prunes identity operations and rewrites eligible transpose convolutions. These passes are enabled by default; each only changes patterns it supports.

Shapes must resolve before code generation. A successful conversion means the compiler could lower this model for the selected settings; it does not prove the firmware produces acceptable outputs on your device.

Select an implementation for each operation

Section titled “Select an implementation for each operation”

Operators check their supported dtypes, layouts, shapes, quantization and platform requirements. Lowering may call a heliaCORE/ns-cmsis-nn kernel, emit inline code, copy data or eliminate a no-op. The operator catalog describes compatibility; an entry is not a promise that every shape and dtype uses an optimized kernel.

Where an operator offers several kernels, applicability predicates and declared priorities select one. Selection is not an on-device search or a measured cost model. Precision explains the additional float and platform restrictions.

The compiler removes unused tensors, interns eligible identical constants and assigns tensors to arenas. It reuses scratch slots when tensor lifetimes do not overlap. Persistent storage carries state across invocations; constant storage supplies weights and other fixed values. The default scratch planner is greedy.

Memory planning gives the module a known layout and arena requirements. Your firmware still has to provide the corresponding memory and linker placement, with space for the rest of the application. Read the memory reports before choosing placement or another planner.

The result includes model/context/tensor code, operator implementation files, constants, headers and the build integration selected by module.type. The supported formats are neuralSPOT, Zephyr, CMake, NSX and CMSIS-Pack. Optional generated tests and reports help inspect and validate the result.

Changing build format does not replace firmware integration or validation. Keep the generated module together with its model, configuration and dependency versions. The generated-files walkthrough shows what to keep and where to start.

  1. InitializePrepare storage, constants and operators
  2. Supply inputs and runUse the model input format and check status
  3. Consume outputsRead results before storage is reused

Use the generated model header for your module’s exact names. With the default module.prefix: aot, the public entry points are aot_model_init() and aot_model_run().

  1. Prepare storage and initialize. Start with a zero-initialized context. In application-owned arena mode, bind every required arena first. Call aot_model_init() and stop on a nonzero status. Initialization resolves tensor pointers, resets persistent state, hydrates staged constants when configured, and initializes operators.
  2. Write inputs and invoke. Use the initialized context’s input views and the model’s dtype, shape and quantization. Fill input bytes before calling aot_model_run(). Check its status before trusting outputs.
  3. Consume outputs, then repeat. Read or copy output data before the next invocation or reinitialization can reuse its storage. Keep the initialized module for successive samples when state should carry. Reinitialize when you deliberately need the model’s reset state.

Follow firmware integration for a complete application path and output validation before using inference results.

The firmware still contains the selected kernels, constants, tensor descriptors and pointer setup, memory arenas, status handling and the generated model’s execution schedule. Staged constants add a copy into their runtime arena during initialization.

The default module.schedule: table runs a generated operator table and supports per-node callbacks. static emits direct calls and omits that callback seam. Both use the same public model API. Neither schedule requires a general model parser on the device; compare the resulting firmware before claiming a size or latency improvement.

Start with defaults, add control for a reason

Section titled “Start with defaults, add control for a reason”

A first conversion already applies default transforms, selects eligible implementations and plans memory. You do not need to configure every pass or tensor to get a module. Enable generated tests when you need an output-verification harness; tests are not enabled by default.

When you need to… Start here Control to explore
Reproduce a conversion Save the model, YAML, package versions and effective settings Configuration precedence
Change a graph rewrite or operator implementation option Identify the affected operation and validate the changed outputs Transforms and operator rules
Fit a memory budget or choose a bank Inspect the arena and tensor residency report Constraints and tensor placement
Move constants into writable memory Account for the source blob, RAM destination and initialization copy Staged constants
Reduce scratch storage Compare layouts for the same model and constraints Scratch planners
Add instrumentation or compare dispatch shapes Keep callbacks with table; measure a static alternative separately Execution settings
Support a custom operation Supply parsing, validation and code generation, with a trusted oracle Custom operators

Transform configuration enables, disables and supplies options to registered passes. It does not reorder them: execution follows registry order. Memory and operator rules likewise change only the matching supported settings; they do not guarantee that every model will fit or become faster.

If you are extending the compiler, continue with Compiler architecture for AIR, the six conversion stages, handler order, operator lifecycle and kernel selection. The generated module reference covers the C integration surface and its distinction from implementation files.