Operators
Use the operator catalog to identify likely compatibility problems, then run conversion to check the actual graph. Support depends on the operator, tensor roles and types, shape, attributes and target; a model name alone cannot answer whether it lowers.
Look up supported operators, tensor data types, kernel dependencies and restrictions before converting your model.
What “supported” means
Section titled “What “supported” means”Three steps must succeed for a remaining node:
- The LiteRT parser produces an AIR operator with a canonical key such as
CONV_2D. - That key resolves to a registered operator class.
- The class accepts the node’s tensor and target conditions and prepares its lowering.
A lowering can be a kernel call, inline arithmetic, a copy, an alias or an accelerator invocation. The operator catalog summarizes the default registry’s types, dependencies and restrictions; a row is not a guarantee of every type/shape pairing or an optimized arithmetic path.
Candidate selection also depends on the target and operator implementation. How heliaAOT works explains the selection stage. The kernel library’s target build remains part of deployment.
Checking a model before you convert
Section titled “Checking a model before you convert”- Compare the model’s operator list and tensor roles with the catalog.
- Read
helia-aot target-info --name <target>for the hardware assumptions. - Convert with the intended configuration and inspect any typed error.
There is no separate per-model coverage-report command. Conversion checks each remaining node before emission, after applicable graph rewrites. Success proves lowering with that configuration; build and output comparison establish the later contracts. No board is needed for this check.
The restrictions every model meets
Section titled “The restrictions every model meets”The input must be a LiteRT flatbuffer and every output extent must resolve at conversion time. A runtime-dependent reshape target cannot be allocated statically. Dynamic signatures that resolve from concrete inputs are a different case from unresolved runtime shapes.
Read requirements for each tensor role. An int16 LSTM cell state does not imply int16 activations, and index/mask tensors can differ from data tensors. The precision guide explains these distinctions.
Restrictions can also involve ranks, constant weights, symmetric zero points, activation functions, broadcasting and signed kernel dimension/context ranges. The catalog is a summary; converter validators remain authoritative. Products can overflow a kernel range even when each dimension fits. Use the error’s node and tensor details to locate the unsupported form.
Stateful models
Section titled “Stateful models”Stateful integration has two distinct contracts. Determine which the exported model uses before writing its streaming loop or golden fixture.
| Operator/path | State contract |
|---|---|
| LSTM and SVDF internal state | Persistent variable tensors survive model_run; initialization resets the internal state. Check each operator’s dtype and shape restrictions. |
| Supported rolled GRU | Explicit initial-state input and final-state output; the caller carries state. Persistent GRU state tensors are rejected. |
| Resource variables | Read/write operations use persistent storage according to the graph’s variable contract. |
The recurrent catalog entries describe supported data types: LSTM and SVDF have
int8 and float paths with role-specific conditions; GRU is FP16, batch size one,
with reset_after=True. Check the complete row for weights, sequence layout,
state and rank constraints.
Where the state lives
Section titled “Where the state lives”Internal variable state is whole-program-live writable storage in a persistent
arena. aot_model_init resets persistent slots and runs operator initialization;
aot_model_run retains the resulting state for the next call. The normal
sequence is:
- Initialize once for a new independent sequence and check the status.
- Copy the next input window and run; check status before consuming output.
- Repeat input/run for subsequent windows without reinitializing.
- Initialize again when deliberately resetting the sequence.
For explicitly exposed state, copy the previous state output to the corresponding
state input, using the model’s I/O order and matching shape/dtype. Use memmove
when buffers may overlap. Reset by supplying the intended initial-state values
before the next independent sequence.
Bindings and generated arena buffers are module-global. Serialize calls and state updates; two context structs are not independent stateful instances. The residency report shows persistent storage, while explicit GRU state must be understood from the model I/O contract. Memory covers arena ownership and sharing.
The rolled GRU
Section titled “The rolled GRU”A Keras export can represent the GRU as a WHILE loop across LiteRT subgraphs.
The default fuse_rolled_gru LiteRT model hook recognizes the supported cell
equations and fuses that form before conversion to AIR. It is distinct from
AIR graph transforms applied later.
The supported GRU has sequence and initial-state inputs, plus sequence and final-state outputs. State tensors must be ordinary graph I/O, not persistent LiteRT variables. Consult the model-level I/O order: it need not match the operator’s local tensor order.
Python registry customizers can replace or remove the hook, but removing it does not make the raw loop lowerable. A different export or corresponding custom parser/lowering is then required.
Carrying state through a test run
Section titled “Carrying state through a test run”For internal persistent state, repeated test iterations retain state without
test.state_feedback. Non-state inputs are refreshed from the same stimulus on
each iteration; this is not a stream of different input windows.
For explicit state I/O, test.state_feedback contains zero-based
[output_position, input_position] pairs. Positions refer to the model’s output
and input lists, not flatbuffer tensor IDs. For a model whose final state is
output 0 and initial state is input 1:
test: enabled: true golden_data: golden.npz num_iterations: 4 state_feedback: - [0, 1]Use this mapping only after checking your model’s order. The converter rejects
out-of-range positions, shape/dtype mismatches and multiple outputs assigned to
the same input. The harness seeds the state input on iteration zero and then
carries output into it with memmove; other inputs are refreshed each iteration.
With a supplied NPZ, output_N must describe the final iteration under that
same stimulus, initial state, count and carry mapping. The converter does not
rerun supplied golden data to repair a mismatched protocol. Without supplied
golden data, its host reference loop uses the configured feedback/count.
For a real sequence of different windows, use an application harness that feeds those windows and checks the required outputs. The generated repeated-stimulus test alone does not establish streaming accuracy. See Testing and automation for the test contract.
When an operator is missing
Section titled “When an operator is missing”First consider a supported upstream export or graph formulation. If custom lowering is needed, supply a LiteRT parser and AOT operator class using the same canonical key. Register them through a Python registry context; Custom operators walks through that process.
When the kernel is not the one you want
Section titled “When the kernel is not the one you want”There is no configuration field selecting an arbitrary kernel name. Selection uses operator conditions, types and platform information. Changing capabilities requires a complete custom target or a separately registered Python platform; capability fields on a registered target name are ignored. Keep those declarations truthful and record the kernel library’s CPU/build flags; see Targets.
A Python customizer can replace an operator class with allow_override=True.
Retain validation and scratch sizing when changing emission, and test the actual
numerics. See replacing a built-in operator.
Rewriting the graph before it is lowered
Section titled “Rewriting the graph before it is lowered”AIR transforms run after loading and before operator resolution. Configure them
in the transforms block described in
Configuring a conversion, or register
a custom transform under its NAME in a Python registry context.
LiteRT model hooks, including rolled-GRU fusion, run earlier over the source model. Choose the extension stage that still contains the structure you need; do not treat a LiteRT multi-subgraph hook and an AIR transform as interchangeable.