Skip to content
heliaAOT
HELIA HUB

Operators

Use the operator catalog to identify likely compatibility problems, then run conversion to check the actual graph. Support depends on the operator, tensor roles and types, shape, attributes and target; a model name alone cannot answer whether it lowers.

Three steps must succeed for a remaining node:

  1. The LiteRT parser produces an AIR operator with a canonical key such as CONV_2D.
  2. That key resolves to a registered operator class.
  3. The class accepts the node’s tensor and target conditions and prepares its lowering.

A lowering can be a kernel call, inline arithmetic, a copy, an alias or an accelerator invocation. The operator catalog summarizes the default registry’s types, dependencies and restrictions; a row is not a guarantee of every type/shape pairing or an optimized arithmetic path.

Candidate selection also depends on the target and operator implementation. How heliaAOT works explains the selection stage. The kernel library’s target build remains part of deployment.

  1. Compare the model’s operator list and tensor roles with the catalog.
  2. Read helia-aot target-info --name <target> for the hardware assumptions.
  3. Convert with the intended configuration and inspect any typed error.

There is no separate per-model coverage-report command. Conversion checks each remaining node before emission, after applicable graph rewrites. Success proves lowering with that configuration; build and output comparison establish the later contracts. No board is needed for this check.

The input must be a LiteRT flatbuffer and every output extent must resolve at conversion time. A runtime-dependent reshape target cannot be allocated statically. Dynamic signatures that resolve from concrete inputs are a different case from unresolved runtime shapes.

Read requirements for each tensor role. An int16 LSTM cell state does not imply int16 activations, and index/mask tensors can differ from data tensors. The precision guide explains these distinctions.

Restrictions can also involve ranks, constant weights, symmetric zero points, activation functions, broadcasting and signed kernel dimension/context ranges. The catalog is a summary; converter validators remain authoritative. Products can overflow a kernel range even when each dimension fits. Use the error’s node and tensor details to locate the unsupported form.

Stateful integration has two distinct contracts. Determine which the exported model uses before writing its streaming loop or golden fixture.

Operator/path State contract
LSTM and SVDF internal state Persistent variable tensors survive model_run; initialization resets the internal state. Check each operator’s dtype and shape restrictions.
Supported rolled GRU Explicit initial-state input and final-state output; the caller carries state. Persistent GRU state tensors are rejected.
Resource variables Read/write operations use persistent storage according to the graph’s variable contract.

The recurrent catalog entries describe supported data types: LSTM and SVDF have int8 and float paths with role-specific conditions; GRU is FP16, batch size one, with reset_after=True. Check the complete row for weights, sequence layout, state and rank constraints.

Internal variable state is whole-program-live writable storage in a persistent arena. aot_model_init resets persistent slots and runs operator initialization; aot_model_run retains the resulting state for the next call. The normal sequence is:

  1. Initialize once for a new independent sequence and check the status.
  2. Copy the next input window and run; check status before consuming output.
  3. Repeat input/run for subsequent windows without reinitializing.
  4. Initialize again when deliberately resetting the sequence.

For explicitly exposed state, copy the previous state output to the corresponding state input, using the model’s I/O order and matching shape/dtype. Use memmove when buffers may overlap. Reset by supplying the intended initial-state values before the next independent sequence.

Bindings and generated arena buffers are module-global. Serialize calls and state updates; two context structs are not independent stateful instances. The residency report shows persistent storage, while explicit GRU state must be understood from the model I/O contract. Memory covers arena ownership and sharing.

A Keras export can represent the GRU as a WHILE loop across LiteRT subgraphs. The default fuse_rolled_gru LiteRT model hook recognizes the supported cell equations and fuses that form before conversion to AIR. It is distinct from AIR graph transforms applied later.

The supported GRU has sequence and initial-state inputs, plus sequence and final-state outputs. State tensors must be ordinary graph I/O, not persistent LiteRT variables. Consult the model-level I/O order: it need not match the operator’s local tensor order.

Python registry customizers can replace or remove the hook, but removing it does not make the raw loop lowerable. A different export or corresponding custom parser/lowering is then required.

For internal persistent state, repeated test iterations retain state without test.state_feedback. Non-state inputs are refreshed from the same stimulus on each iteration; this is not a stream of different input windows.

For explicit state I/O, test.state_feedback contains zero-based [output_position, input_position] pairs. Positions refer to the model’s output and input lists, not flatbuffer tensor IDs. For a model whose final state is output 0 and initial state is input 1:

test:
enabled: true
golden_data: golden.npz
num_iterations: 4
state_feedback:
- [0, 1]

Use this mapping only after checking your model’s order. The converter rejects out-of-range positions, shape/dtype mismatches and multiple outputs assigned to the same input. The harness seeds the state input on iteration zero and then carries output into it with memmove; other inputs are refreshed each iteration.

With a supplied NPZ, output_N must describe the final iteration under that same stimulus, initial state, count and carry mapping. The converter does not rerun supplied golden data to repair a mismatched protocol. Without supplied golden data, its host reference loop uses the configured feedback/count.

For a real sequence of different windows, use an application harness that feeds those windows and checks the required outputs. The generated repeated-stimulus test alone does not establish streaming accuracy. See Testing and automation for the test contract.

First consider a supported upstream export or graph formulation. If custom lowering is needed, supply a LiteRT parser and AOT operator class using the same canonical key. Register them through a Python registry context; Custom operators walks through that process.

There is no configuration field selecting an arbitrary kernel name. Selection uses operator conditions, types and platform information. Changing capabilities requires a complete custom target or a separately registered Python platform; capability fields on a registered target name are ignored. Keep those declarations truthful and record the kernel library’s CPU/build flags; see Targets.

A Python customizer can replace an operator class with allow_override=True. Retain validation and scratch sizing when changing emission, and test the actual numerics. See replacing a built-in operator.

AIR transforms run after loading and before operator resolution. Configure them in the transforms block described in Configuring a conversion, or register a custom transform under its NAME in a Python registry context.

LiteRT model hooks, including rolled-GRU fusion, run earlier over the source model. Choose the extension stage that still contains the structure you need; do not treat a LiteRT multi-subgraph hook and an AIR transform as interchangeable.