# Compiler architecture

This page is for developers extending parsing, transforms or code generation. For the deployment mental model, start with [How heliaAOT works](https://ambiqai.github.io/helia-aot/guide/how-it-works/). You do not need to implement handlers or understand AIR to convert a supported model.

## Inside the compiler

`AotConverter.convert()` loads a model, transforms its graph, resolves operators and handlers, plans storage, emits files and exports the module. `--verbose 2` exposes conversion details. The implementation lives in `helia_aot/converter.py`; generated [Python API reference](https://ambiqai.github.io/helia-aot/reference/api/helia_aot/) supplies the exact entry points.

| Stage | Main responsibility | Failure to investigate here |
| --- | --- | --- |
| Load | Parse the selected LiteRT subgraph into AIR | Missing model, unsupported input format or parser |
| Transform | Apply graph rewrites, then propagate shapes strictly | Invalid transform name or unresolved shape |
| Resolve | Validate and lower operators; prepare handlers | Unsupported operator contract, dtype, layout or platform |
| Plan | Optimize resolved tensors, assign arenas and prepare render data | Memory constraints or alignment cannot be satisfied |
| Emit | Render module and operator templates | Missing template data or invalid extension implementation |
| Export | Write directory, ZIP or CMSIS-Pack output | Existing destination or incompatible packaging choice |

## Load

The front end accepts LiteRT flatbuffers. `model.type`, when supplied, selects the parser; otherwise the file suffix is used. Recognized types are `tflite` and `litert`. Renaming another format does not convert it.

Registered model hooks can rewrite the flatbuffer before parsing; the built-in rolled-GRU hook recognizes supported loop patterns. The parser walks the selected subgraph and dispatches through an explicit `RegistryContext` using canonical operator keys.

AIR represents tensors with shape, dtype, quantization and storage kind, and operations with input/output tensor IDs and typed options. Recurrent state contracts depend on the operator: explicit GRU state I/O differs from an LSTM's internal persistent variable tensors. See [Operators](https://ambiqai.github.io/helia-aot/guide/operators/).

## Transform

Default transforms run in registry order:

1. `FOLD_STATIC_SHAPE_EXPRESSIONS`
2. `DEPTHWISE_TO_CONV`
3. `PRUNE_IDENTITY_OPS`
4. `TRANSPOSE_REVERSE_CONV`

An empty `transforms` list enables all registered transforms. Configuration entries replace the enabled/options settings for a matching name; their order does not reorder execution. A `name: "*"` entry sets the default enabled state. Supply options on named entries: wildcard options are not copied into every pass.

Strict shape propagation runs before graph transforms, resolving shapes and static shape-expression values together. `FOLD_STATIC_SHAPE_EXPRESSIONS` then replaces evaluated expressions with constants. Each pass acts only on its eligible patterns; unresolved dimensions fail before code generation. When adding a transform, validate semantic equivalence on representative models rather than treating a smaller graph as proof of correctness.

## Resolve

### Handlers

Handlers run in this order: `ModuleHandler`, `OperatorHandler`, `TensorHandler`, `ModelHandler`, `DocHandler`, then `TestHandler` when `test.enabled` is true.

| Handler | Owns |
| --- | --- |
| Module | Common headers, build-format files and kernel-library dependency requirements |
| Operator | Per-node operator classes, selected implementation and templates |
| Tensor | Tensor descriptors, arenas and constant data |
| Model | Initialization, execution schedule and public model entry points |
| Documentation | Generated README/licence and optional offline HTML documentation |
| Test | Optional generated input/output validation harness |

The converter invokes each handler's `resolve()`, later its `plan()`, then its `emit()`. Earlier handlers' decisions are available to later ones. The resolved model is optimized after resolution, before memory planning.

### The operator lifecycle

An operator validates ranks, layouts, dtypes, quantization and target requirements; resolves implementation and temporary-storage needs; declares planned storage; computes template values; and emits code. The base class and each subclass determine which hooks implement those steps. The template environment is strict: missing required values fail emission.

Runtime initialization is separate from this compiler lifecycle. Operators without an initialization requirement can declare `has_init = False`; generated model code omits their init call while retaining the table-mode callback contract. See the maintained [custom-operator example](https://ambiqai.github.io/helia-aot/examples/custom-operator/) before adding parser/emitter classes.

### Kernel selection

Selection is local to the operator. Operators such as `CONV_2D`, `DEPTHWISE_CONV_2D` and `TRANSPOSE_CONV` maintain candidate tables in their `kernels.py` modules. Candidates declare a C symbol, capability requirements, shape/attribute predicates, priority and scratch requirements.

The resolver filters candidates against the resolved platform and operation, then chooses the applicable candidate with highest declared priority. No applicable candidate means conversion fails. A higher priority is an implementation policy, not a latency estimate or a benchmark result. Many operators have a single implementation and no candidate table.

### Dtype gates and the float switch

Dtype membership is only the first compatibility gate. An implementation can impose additional shape, quantization, tensor-role and platform restrictions. Some lowering emits inline or reference code; a supported dtype does not imply a native optimized kernel.

Operators that bind native float kernels declare the corresponding library requirement. FP16 additionally requires an admitted platform: explicit FP16 capability with a compatible core, or the supported Cortex-M55 fallback. Storage dtype alone is not evidence of half-precision arithmetic. See [Precision](https://ambiqai.github.io/helia-aot/guide/precision/) for build flags and restrictions.

### The ns-cmsis-nn version floor

Operators declare minimum kernel-library versions. The module combines those requirements, emits a version check in the common header, and records the dependency in CMSIS-Pack metadata. Use the generated files' requirement when integrating; a library version check does not itself validate every compiler, ABI or link setting.

## Plan

The resolved-model optimizer can intern identical constants; the converter prunes unused tensors before calling the selected planner. `memory.planner` defaults to `greedy`; `greedy_by_size` and `hill_climb` are experimental alternatives.

The plan binds each tensor to a region and offset. Scratch slots can be reused across non-overlapping lifetimes. Persistent storage survives ordinary invocations and is reset during initialization. Constants are read in place or copied from a source blob into a staged runtime arena. These roles describe storage; they do not allocate an independent state per C context struct.

A render plan groups arenas and packed data for templates. Plan and tensor-layout hashes identify layout changes; they are not model-content hashes or numerical correctness checks. [Memory planning](https://ambiqai.github.io/helia-aot/guide/memory/) and [reports](https://ambiqai.github.io/helia-aot/guide/memory-reports/) explain how to inspect the result.

## Emit

Handlers render into the converter's working directory. Typical output includes model, context, tensor, constant and operator sources; public and implementation headers; build glue; documentation; and optional test code. The exact inventory depends on the model and settings. Descriptor-based operators may share a kernel body instead of emitting it for every node.

`module.prefix` namespaces generated symbols. `module.schedule: table` uses a function-pointer table and per-node callbacks. `static` emits direct calls without that callback seam. Both retain initialization/status handling and the same public model API.

Use [Generated module](https://ambiqai.github.io/helia-aot/reference/module/) to distinguish the supported integration surface from headers and symbols that merely appear in a reference fixture. Applications should not depend on per-node file names or private context fields.

## Export

For a directory destination, the module is written under `module.path/module.name`. A `.zip` destination contains the module directory. A `.pack` destination requires `module.type: cmsis_pack` and puts the pack contents, including its manifest, at the archive root. Existing output requires `force` to overwrite; use distinct paths while comparing configurations.

CMSIS-Pack export honors `SOURCE_DATE_EPOCH` for fixed archive timestamps. Ordinary ZIP export in this revision uses the standard archive writer; do not infer byte-for-byte reproducibility for all outputs from the pack setting.

The final console summary names the output and arena requirements. `memory.dump_residency_json` writes the optional machine-readable memory report. Preserve model/configuration/tool versions alongside outputs when you need a reproducible conversion record.
