Skip to content
heliaAOT
HELIA HUB

Compiler architecture

This page is for developers extending parsing, transforms or code generation. For the deployment mental model, start with How heliaAOT works. You do not need to implement handlers or understand AIR to convert a supported model.

AotConverter.convert() loads a model, transforms its graph, resolves operators and handlers, plans storage, emits files and exports the module. --verbose 2 exposes conversion details. The implementation lives in helia_aot/converter.py; generated Python API reference supplies the exact entry points.

Stage Main responsibility Failure to investigate here
Load Parse the selected LiteRT subgraph into AIR Missing model, unsupported input format or parser
Transform Apply graph rewrites, then propagate shapes strictly Invalid transform name or unresolved shape
Resolve Validate and lower operators; prepare handlers Unsupported operator contract, dtype, layout or platform
Plan Optimize resolved tensors, assign arenas and prepare render data Memory constraints or alignment cannot be satisfied
Emit Render module and operator templates Missing template data or invalid extension implementation
Export Write directory, ZIP or CMSIS-Pack output Existing destination or incompatible packaging choice

The front end accepts LiteRT flatbuffers. model.type, when supplied, selects the parser; otherwise the file suffix is used. Recognized types are tflite and litert. Renaming another format does not convert it.

Registered model hooks can rewrite the flatbuffer before parsing; the built-in rolled-GRU hook recognizes supported loop patterns. The parser walks the selected subgraph and dispatches through an explicit RegistryContext using canonical operator keys.

AIR represents tensors with shape, dtype, quantization and storage kind, and operations with input/output tensor IDs and typed options. Recurrent state contracts depend on the operator: explicit GRU state I/O differs from an LSTM’s internal persistent variable tensors. See Operators.

Default transforms run in registry order:

  1. FOLD_STATIC_SHAPE_EXPRESSIONS
  2. DEPTHWISE_TO_CONV
  3. PRUNE_IDENTITY_OPS
  4. TRANSPOSE_REVERSE_CONV

An empty transforms list enables all registered transforms. Configuration entries replace the enabled/options settings for a matching name; their order does not reorder execution. A name: "*" entry sets the default enabled state. Supply options on named entries: wildcard options are not copied into every pass.

Strict shape propagation runs before graph transforms, resolving shapes and static shape-expression values together. FOLD_STATIC_SHAPE_EXPRESSIONS then replaces evaluated expressions with constants. Each pass acts only on its eligible patterns; unresolved dimensions fail before code generation. When adding a transform, validate semantic equivalence on representative models rather than treating a smaller graph as proof of correctness.

Handlers run in this order: ModuleHandler, OperatorHandler, TensorHandler, ModelHandler, DocHandler, then TestHandler when test.enabled is true.

Handler Owns
Module Common headers, build-format files and kernel-library dependency requirements
Operator Per-node operator classes, selected implementation and templates
Tensor Tensor descriptors, arenas and constant data
Model Initialization, execution schedule and public model entry points
Documentation Generated README/licence and optional offline HTML documentation
Test Optional generated input/output validation harness

The converter invokes each handler’s resolve(), later its plan(), then its emit(). Earlier handlers’ decisions are available to later ones. The resolved model is optimized after resolution, before memory planning.

An operator validates ranks, layouts, dtypes, quantization and target requirements; resolves implementation and temporary-storage needs; declares planned storage; computes template values; and emits code. The base class and each subclass determine which hooks implement those steps. The template environment is strict: missing required values fail emission.

Runtime initialization is separate from this compiler lifecycle. Operators without an initialization requirement can declare has_init = False; generated model code omits their init call while retaining the table-mode callback contract. See the maintained custom-operator example before adding parser/emitter classes.

Selection is local to the operator. Operators such as CONV_2D, DEPTHWISE_CONV_2D and TRANSPOSE_CONV maintain candidate tables in their kernels.py modules. Candidates declare a C symbol, capability requirements, shape/attribute predicates, priority and scratch requirements.

The resolver filters candidates against the resolved platform and operation, then chooses the applicable candidate with highest declared priority. No applicable candidate means conversion fails. A higher priority is an implementation policy, not a latency estimate or a benchmark result. Many operators have a single implementation and no candidate table.

Dtype membership is only the first compatibility gate. An implementation can impose additional shape, quantization, tensor-role and platform restrictions. Some lowering emits inline or reference code; a supported dtype does not imply a native optimized kernel.

Operators that bind native float kernels declare the corresponding library requirement. FP16 additionally requires an admitted platform: explicit FP16 capability with a compatible core, or the supported Cortex-M55 fallback. Storage dtype alone is not evidence of half-precision arithmetic. See Precision for build flags and restrictions.

Operators declare minimum kernel-library versions. The module combines those requirements, emits a version check in the common header, and records the dependency in CMSIS-Pack metadata. Use the generated files’ requirement when integrating; a library version check does not itself validate every compiler, ABI or link setting.

The resolved-model optimizer can intern identical constants; the converter prunes unused tensors before calling the selected planner. memory.planner defaults to greedy; greedy_by_size and hill_climb are experimental alternatives.

The plan binds each tensor to a region and offset. Scratch slots can be reused across non-overlapping lifetimes. Persistent storage survives ordinary invocations and is reset during initialization. Constants are read in place or copied from a source blob into a staged runtime arena. These roles describe storage; they do not allocate an independent state per C context struct.

A render plan groups arenas and packed data for templates. Plan and tensor-layout hashes identify layout changes; they are not model-content hashes or numerical correctness checks. Memory planning and reports explain how to inspect the result.

Handlers render into the converter’s working directory. Typical output includes model, context, tensor, constant and operator sources; public and implementation headers; build glue; documentation; and optional test code. The exact inventory depends on the model and settings. Descriptor-based operators may share a kernel body instead of emitting it for every node.

module.prefix namespaces generated symbols. module.schedule: table uses a function-pointer table and per-node callbacks. static emits direct calls without that callback seam. Both retain initialization/status handling and the same public model API.

Use Generated module to distinguish the supported integration surface from headers and symbols that merely appear in a reference fixture. Applications should not depend on per-node file names or private context fields.

For a directory destination, the module is written under module.path/module.name. A .zip destination contains the module directory. A .pack destination requires module.type: cmsis_pack and puts the pack contents, including its manifest, at the archive root. Existing output requires force to overwrite; use distinct paths while comparing configurations.

CMSIS-Pack export honors SOURCE_DATE_EPOCH for fixed archive timestamps. Ordinary ZIP export in this revision uses the standard archive writer; do not infer byte-for-byte reproducibility for all outputs from the pack setting.

The final console summary names the output and arena requirements. memory.dump_residency_json writes the optional machine-readable memory report. Preserve model/configuration/tool versions alongside outputs when you need a reproducible conversion record.