Testing and automation
Use the generated test harness to compare a module with trusted outputs, then repeat conversion, build and validation in your pipeline. This guide covers golden fixtures, tolerance and state handling, followed by exit codes and reports for automation.
Validate it on target is the walkthrough for a first module. This page is the detail behind it.
Start with one checked inference
Section titled “Start with one checked inference”Use a model that already converts and a target build that can call its generated
C API. Add test.enabled: true to the conversion configuration. If you have
trusted expected outputs, supply a golden archive; otherwise the converter
needs a host interpreter for the model.
helia-aot convert --path convert.yaml --test.enabledBuild the generated test source with the module in your target project, then
call aot_test_case_init() once and aot_test_case_run(). With the default
prefix aot, both functions return 0 on success. Check both results; an
initialization failure must stop the harness. The first validation
walkthrough supplies the integration path.
With output verification enabled, a successful run means the supplied stimulus produced outputs within the configured tolerance of the chosen oracle. It does not establish model accuracy over a dataset or correctness for every input. Keep the oracle, tolerance, model, configuration and target build identity with the result before changing conversion controls. This shortest path uses the default internally allocated arenas; external buffers have the additional harness requirements below.
The generated test case
Section titled “The generated test case”Set test.enabled and the conversion emits aot_test_case.c and its header
alongside the module, with the stimulus and the expected outputs baked in as
constant arrays. The test needs no file system on the board and reads nothing
at run time.
test: enabled: true golden_data: golden.npz tolerance: 1.0 num_iterations: 1 skip_verification: false| Field | Default | What it does |
|---|---|---|
test.enabled |
false |
Emits the test case source and header. |
test.golden_data |
none | Path to an .npz holding the stimulus and the expected outputs. |
test.tolerance |
1.0 |
Absolute per-element difference allowed in the output dtype’s units; see tolerance. |
test.num_iterations |
1 |
How many times the stimulus is run before the outputs are checked. |
test.state_feedback |
empty | (output_index, input_index) pairs carried between iterations. |
test.skip_verification |
false |
Runs the model but does not compare the outputs. |
Configuration is the generated list of every field and its default.
The golden file
Section titled “The golden file”test.golden_data points at a NumPy .npz archive. It is loaded with pickling
disabled, so it must hold plain numeric arrays and nothing else.
| Key | Holds | Required |
|---|---|---|
input_0 |
The stimulus for the model’s first input | One key per model input |
input_1, input_2, … |
The stimulus for each further input, in model input order | Yes |
output_0 |
The expected value of the model’s first output | One key per model output |
output_1, output_2, … |
The expected value of each further output, in model output order | Yes |
The index in the key is the position of the tensor in the model’s input or output list, not the tensor’s own identifier. Each array is written into that tensor, so its shape and element type are the tensor’s shape and element type.
Provide arrays with exactly the model tensor’s shape and dtype. The loader checks required keys, but do not rely on it to diagnose every shape or dtype mistake in an archive you supply.
A key the conversion expects and does not find fails the conversion with a
GoldenDataKeyError naming the key and the file. There is no fallback to
another source of truth: the archive supplies the expected values even when
skip_verification disables runtime comparison.
For the complete first-deployment path, use the supplied, hash-checked fixture
in First conversion.
For your own model, first obtain stimulus and expected NumPy arrays from
an evaluation you trust, already in the model’s input/output shapes and dtypes.
Then this NPZ-writing excerpt saves a single-input, single-output case:
import numpy as np
np.savez( "golden.npz", input_0=stimulus, output_0=expected,)This excerpt does not generate inputs or run a reference model. Casting a floating array to int8 is not quantization; apply the model’s scale and zero point in your evaluation pipeline and check the stored values before saving.
The three oracles
Section titled “The three oracles”Expected data can come from your archive, a host interpreter, or a placeholder
when an interpreter is unavailable. Separately, skip_verification decides
whether the generated test compares any outputs. These combinations make
different claims:
golden_data |
skip_verification |
What the expected outputs are | What a pass means |
|---|---|---|---|
| Set | false |
The output_N arrays from your file |
The module matched your expected outputs within tolerance |
| Set | true |
The file is loaded, but outputs are not compared | The supplied stimulus ran without a reported runtime error |
| Not set | false |
A host interpreter run over deterministic stimulus derived from the model name and tensor | The module matched that reference execution within tolerance |
| Not set | true |
Interpreter outputs if available; otherwise zero placeholders; neither is compared | The stimulus ran without a reported runtime error |
The no-file, verification-enabled case is a self-consistency check between the generated C and a host interpreter, not a check against data you picked. It catches a lowering that changed behaviour. It cannot catch a model that was wrong before heliaAOT saw it.
Both rows with skip_verification: true are smoke tests. The converter still
attempts to create a host interpreter when no golden file is supplied, and may
run it if creation succeeds. The flag permits conversion without one; it does
not mean that expected outputs will be checked later on the device. The
printed Test passed! marker also appears after a successful smoke run; require
verification-enabled configuration and per-output comparison records before
reporting a numerical match.
Choosing a tolerance
Section titled “Choosing a tolerance”The comparison is absolute and per element, in the stored output values.
For integer output it compares integer differences and casts test.tolerance
to an integer: 1.5 therefore permits a difference of 1. It does not compare
dequantized values or relative error. For floating output it compares the
absolute difference using float arithmetic and a float tolerance.
The default 1.0 is not an accuracy recommendation for every model. Choose a
non-negative, finite tolerance from your numerical requirements and trusted
reference, then investigate any mismatch before changing it. A tolerance
suitable for quantized integer values may be far too loose for floating outputs.
For floating outputs, an expected NaN matches an actual NaN, and an expected infinity matches only an infinity of the same sign. When the expected value is finite, an actual NaN or infinity fails. This policy checks agreement with the oracle; an oracle containing NaNs still needs review for your application.
Iterations and state
Section titled “Iterations and state”num_iterations runs the stimulus through the model repeatedly before the
outputs are checked once, at the end. On a stateless model this only exercises
the run path; the extra passes change nothing.
Internal persistent state can evolve across invocations without a feedback
setting. For a model that exposes state as inputs and outputs, state_feedback
names the outputs to copy back into inputs between iterations. Each pair is
(output_index, input_index) in model I/O order, must have matching shape and
dtype, and may drive an input only once. For example, state_feedback: [[1, 1]]
feeds the second output into the second input when those tensors form a
compatible state pair.
The host interpreter carries that feedback through the same number of
iterations. When you supply an archive instead, its output_N arrays must
already contain the expected final-iteration outputs for that stimulus
and initial state. The test checks outputs once, after the last iteration.
Operators
covers the pairs and how the generated code seeds each kind of input.
To repeat the test from fresh persistent state, initialize it again before
calling test_case_run. Within one run, num_iterations intentionally carries
state without reinitializing the model.
External-arena tests
Section titled “External-arena tests”With memory.allocate_arenas: false, your application must load and bind cold
constant sidecars. The generated test harness does not perform that load. It
returns -1 for external-arena configurations with non-writable cold constant
arenas. For cold constants in writable memory it can allocate a buffer, but
that buffer does not contain the sidecar weights and cannot establish correct
inference. Supply a harness that loads the emitted __blob.bin bytes, binds
every region, initializes the model and checks outputs against the same oracle.
See Caller-supplied arenas.
The two functions
Section titled “The two functions”int32_t aot_test_case_init(void);int32_t aot_test_case_run(void);Both return int32_t and both return 0 on success, so a harness can treat
either non-zero return as a failure. They fail for different reasons, and the
difference is worth keeping in your logs: a non-zero return from
aot_test_case_init is an initialization failure and nothing has run yet,
while a non-zero return from aot_test_case_run means either an operator
returned an error status or at least one output element exceeded the tolerance.
Automating conversions
Section titled “Automating conversions”Exit codes
Section titled “Exit codes”Check the conversion exit code before building or publishing its output:
| Code | Meaning |
|---|---|
| 0 | The module was written. |
| 1 | A typed conversion failure, or an unhandled exception; inspect the error output. |
| 2 | The arguments did not validate, with one message per offending field. |
| 130 | Interrupted. |
Typed conversion failures have a message and may include a remediation hint. Some input failures still raise standard exceptions, including a missing golden archive or unknown transform name. Preserve the traceback and check Troubleshooting before reporting a defect.
Converting into a directory that already holds that module is refused rather
than silently overwritten, so a pipeline either cleans its output directory or
passes --force deliberately.
The machine-readable report
Section titled “The machine-readable report”Set memory.dump_residency_json and the conversion writes
<prefix>_residency.json next to the module. That file, not the console
output, is what a pipeline should read.
helia-aot convert --path kws.yaml --memory.dump-residency-jsonIt carries a schema version, the arena envelope for each role, every tensor
with its role, memory, offset and size, and two hashes: plan_hash over the
arena envelope, and tensor_layout_hash over per-tensor placement. Gate on it
the way you would gate on a linker map: fail the build when an arena grows past
what the board has, and treat a hash change as a change that needs reviewing.
Memory reads a full report
field by field.
Comparing two conversions
Section titled “Comparing two conversions”Compare the generated source as well as the residency reports. For example, keep baseline and candidate conversions in separate directories:
diff -ru out-baseline/my_model out-candidate/my_modelA difference exits with status 1; an execution error exits with status 2. Review the changed files alongside model/configuration versions before accepting a candidate. A source diff explains code changes, not their performance impact.
That is also the pattern this repository uses on itself. A set of generated files is committed as baselines, the unit suite regenerates them and compares byte for byte, and a change that alters codegen fails until the baselines are updated deliberately. Applied to your own project, it means committing a converted module for one representative model and letting the diff be the review.
What this repository’s CI covers
Section titled “What this repository’s CI covers”The checks below run on heliaAOT itself. They are not checks on your model, and a green run covers those tested cases; it does not validate your conversion. They are worth knowing because they set the floor of what is known to work.
| Job | What it does |
|---|---|
lint-unit |
Lint and the unit suite, on Python 3.11 and 3.14 |
e2e |
Full conversions built and executed for apollo4p_blue_kbr_evb and apollo510_evb on the Corstone-300 fixed virtual platform |
zephyr-build |
Compiles a generated Zephyr module, as a compile check rather than a run |
Your own pipeline is the mirror image of that: convert, build, and run the generated test case on the board or on a model of it, then keep the residency report and the emitted tree as artifacts so the next run has something to be compared against.