# How heliaAOT works

heliaAOT turns an exported LiteRT model into a model-specific C module. Your development machine does the graph processing, kernel selection and memory planning. Your firmware initializes the generated module, supplies input data and calls it to run inference.

You still build and link firmware for your board. The generated module uses kernel-library code and a small amount of generated execution support; it does not need a LiteRT interpreter to parse and plan the model on the device.

## From model and settings to a module

1. Your model and settings: 01 · Input
   - [LiteRT model](https://ambiqai.github.io/helia-aot/getting-started/targets/): Tensors, operators and metadata
   - [Configuration](https://ambiqai.github.io/helia-aot/guide/configuring/): Target, output format and optional controls
2. heliaAOT: 02 · Compile on your development machine
   - [Graph processing](https://ambiqai.github.io/helia-aot/guide/how-it-works/#analyze-and-transform-the-model): Analyze shapes and transform supported patterns
   - [Target and kernels](https://ambiqai.github.io/helia-aot/guide/how-it-works/#select-an-implementation-for-each-operation): Match operations to eligible implementations
   - [Memory planning](https://ambiqai.github.io/helia-aot/guide/how-it-works/#lay-out-memory-before-execution): Assign arenas, reuse scratch and place tensors
3. Model code and artifacts: 03 · Output
   - Operator code: Generated C and selected kernel calls
   - [Memory arenas](https://ambiqai.github.io/helia-aot/guide/memory/): Constants, persistent state and scratch
   - [Execution and APIs](https://ambiqai.github.io/helia-aot/reference/module/): Initialization and a generated model schedule
   - [Packaging](https://ambiqai.github.io/helia-aot/guide/memory-placement/): Embedded constants or configured sidecar blobs
   - Documentation: Module README and optional HTML or memory reports
   - [Build integration](https://ambiqai.github.io/helia-aot/getting-started/integrate/): CMake, Zephyr, neuralSPOT, NSX or CMSIS-Pack

This is a conceptual build-time view. The output is a module ready for firmware integration, not a board executable: your application still compiles and links it with the selected kernels and platform support. Useful defaults mean you do not need to specify every deployment setting.

Training and model export happen upstream. The compiler accepts LiteRT flatbuffers (`.tflite` or `.litert`); it does not train a network or automatically convert arbitrary models between integer and float precision. Start with the [first-conversion walkthrough](https://ambiqai.github.io/helia-aot/getting-started/convert/) for a supplied model and configuration.

## Before deployment

The conversion makes decisions that would otherwise need runtime machinery or hand-written integration work.

1. [Understand the model](https://ambiqai.github.io/helia-aot/guide/how-it-works/#analyze-and-transform-the-model): Operators, shapes, precision and support
2. [Optimize and select kernels](https://ambiqai.github.io/helia-aot/guide/how-it-works/#select-an-implementation-for-each-operation): Graph rewrites and eligible target implementations
3. [Plan memory](https://ambiqai.github.io/helia-aot/guide/how-it-works/#lay-out-memory-before-execution): Constants, persistent state, scratch and placement
4. [Generate deployment](https://ambiqai.github.io/helia-aot/guide/how-it-works/#write-a-module-you-can-build): Model code, integration and enabled reports or tests

These blocks expand the heliaAOT box above into user-relevant decisions, rather than specifying an exact compiler pass order. The [compiler architecture](https://ambiqai.github.io/helia-aot/guide/compiler-architecture/) gives that detailed sequence. The controls table below links each decision to optional settings.

### Analyze and transform the model

The compiler reads the selected subgraph, applies registered model hooks and runs graph transforms. The built-in transform pipeline folds static shape expressions, rewrites eligible depthwise convolutions, prunes identity operations and rewrites eligible transpose convolutions. These passes are enabled by default; each only changes patterns it supports.

Shapes must resolve before code generation. A successful conversion means the compiler could lower this model for the selected settings; it does not prove the firmware produces acceptable outputs on your device.

### Select an implementation for each operation

Operators check their supported dtypes, layouts, shapes, quantization and platform requirements. Lowering may call a heliaCORE/ns-cmsis-nn kernel, emit inline code, copy data or eliminate a no-op. The [operator catalog](https://ambiqai.github.io/helia-aot/reference/operators/) describes compatibility; an entry is not a promise that every shape and dtype uses an optimized kernel.

Where an operator offers several kernels, applicability predicates and declared priorities select one. Selection is not an on-device search or a measured cost model. [Precision](https://ambiqai.github.io/helia-aot/guide/precision/) explains the additional float and platform restrictions.

### Lay out memory before execution

The compiler removes unused tensors, interns eligible identical constants and assigns tensors to arenas. It reuses scratch slots when tensor lifetimes do not overlap. Persistent storage carries state across invocations; constant storage supplies weights and other fixed values. The default scratch planner is `greedy`.

Memory planning gives the module a known layout and arena requirements. Your firmware still has to provide the corresponding memory and linker placement, with space for the rest of the application. Read the [memory reports](https://ambiqai.github.io/helia-aot/guide/memory-reports/) before choosing placement or another planner.

### Write a module you can build

The result includes model/context/tensor code, operator implementation files, constants, headers and the build integration selected by `module.type`. The supported formats are neuralSPOT, Zephyr, CMake, NSX and CMSIS-Pack. Optional generated tests and reports help inspect and validate the result.

Changing build format does not replace firmware integration or validation. Keep the generated module together with its model, configuration and dependency versions. The [generated-files walkthrough](https://ambiqai.github.io/helia-aot/getting-started/module/) shows what to keep and where to start.

## On the device

1. Initialize: Prepare storage, constants and operators
2. Supply inputs and run: Use the model input format and check status
3. Consume outputs: Read results before storage is reused

Use the generated model header for your module's exact names. With the default `module.prefix: aot`, the public entry points are `aot_model_init()` and `aot_model_run()`.

1. **Prepare storage and initialize.** Start with a zero-initialized context. In application-owned arena mode, bind every required arena first. Call `aot_model_init()` and stop on a nonzero status. Initialization resolves tensor pointers, resets persistent state, hydrates staged constants when configured, and initializes operators.
2. **Write inputs and invoke.** Use the initialized context's input views and the model's dtype, shape and quantization. Fill input bytes before calling `aot_model_run()`. Check its status before trusting outputs.
3. **Consume outputs, then repeat.** Read or copy output data before the next invocation or reinitialization can reuse its storage. Keep the initialized module for successive samples when state should carry. Reinitialize when you deliberately need the model's reset state.

Follow [firmware integration](https://ambiqai.github.io/helia-aot/getting-started/integrate/) for a complete application path and [output validation](https://ambiqai.github.io/helia-aot/getting-started/validate/) before using inference results.

### What remains at runtime

The firmware still contains the selected kernels, constants, tensor descriptors and pointer setup, memory arenas, status handling and the generated model's execution schedule. Staged constants add a copy into their runtime arena during initialization.

The default `module.schedule: table` runs a generated operator table and supports per-node callbacks. `static` emits direct calls and omits that callback seam. Both use the same public model API. Neither schedule requires a general model parser on the device; compare the resulting firmware before claiming a size or latency improvement.

:::caution[One module is one shared execution state]
A second context struct does not allocate independent model arenas. Arena bindings and hydration state are module-global. Serialize use of a module, and keep its buffers valid through execution. Separately generated modules need distinct prefixes; sharing their scratch memory requires non-overlapping execution. See [memory ownership](https://ambiqai.github.io/helia-aot/guide/memory-placement/).
:::

## Start with defaults, add control for a reason

A first conversion already applies default transforms, selects eligible implementations and plans memory. You do not need to configure every pass or tensor to get a module. Enable generated tests when you need an output-verification harness; tests are not enabled by default.

| When you need to… | Start here | Control to explore |
| --- | --- | --- |
| Reproduce a conversion | Save the model, YAML, package versions and effective settings | [Configuration precedence](https://ambiqai.github.io/helia-aot/guide/configuring/) |
| Change a graph rewrite or operator implementation option | Identify the affected operation and validate the changed outputs | [Transforms and operator rules](https://ambiqai.github.io/helia-aot/guide/configuring/#transforms) |
| Fit a memory budget or choose a bank | Inspect the arena and tensor residency report | [Constraints and tensor placement](https://ambiqai.github.io/helia-aot/guide/memory-placement/) |
| Move constants into writable memory | Account for the source blob, RAM destination and initialization copy | [Staged constants](https://ambiqai.github.io/helia-aot/guide/memory-placement/) |
| Reduce scratch storage | Compare layouts for the same model and constraints | [Scratch planners](https://ambiqai.github.io/helia-aot/guide/memory-planners/) |
| Add instrumentation or compare dispatch shapes | Keep callbacks with `table`; measure a `static` alternative separately | [Execution settings](https://ambiqai.github.io/helia-aot/guide/configuring/#schedule-mode) |
| Support a custom operation | Supply parsing, validation and code generation, with a trusted oracle | [Custom operators](https://ambiqai.github.io/helia-aot/guide/custom-operators/) |

Transform configuration enables, disables and supplies options to registered passes. It does **not** reorder them: execution follows registry order. Memory and operator rules likewise change only the matching supported settings; they do not guarantee that every model will fit or become faster.

## Compiler details

If you are extending the compiler, continue with [Compiler architecture](https://ambiqai.github.io/helia-aot/guide/compiler-architecture/) for AIR, the six conversion stages, handler order, operator lifecycle and kernel selection. The [generated module reference](https://ambiqai.github.io/helia-aot/reference/module/) covers the C integration surface and its distinction from implementation files.
