Skip to content
heliaAOT
HELIA HUB

Testing and automation

Use the generated test harness to compare a module with trusted outputs, then repeat conversion, build and validation in your pipeline. This guide covers golden fixtures, tolerance and state handling, followed by exit codes and reports for automation.

Validate it on target is the walkthrough for a first module. This page is the detail behind it.

Use a model that already converts and a target build that can call its generated C API. Add test.enabled: true to the conversion configuration. If you have trusted expected outputs, supply a golden archive; otherwise the converter needs a host interpreter for the model.

Terminal window
helia-aot convert --path convert.yaml --test.enabled

Build the generated test source with the module in your target project, then call aot_test_case_init() once and aot_test_case_run(). With the default prefix aot, both functions return 0 on success. Check both results; an initialization failure must stop the harness. The first validation walkthrough supplies the integration path.

With output verification enabled, a successful run means the supplied stimulus produced outputs within the configured tolerance of the chosen oracle. It does not establish model accuracy over a dataset or correctness for every input. Keep the oracle, tolerance, model, configuration and target build identity with the result before changing conversion controls. This shortest path uses the default internally allocated arenas; external buffers have the additional harness requirements below.

Set test.enabled and the conversion emits aot_test_case.c and its header alongside the module, with the stimulus and the expected outputs baked in as constant arrays. The test needs no file system on the board and reads nothing at run time.

kws.yaml
test:
enabled: true
golden_data: golden.npz
tolerance: 1.0
num_iterations: 1
skip_verification: false
Field Default What it does
test.enabled false Emits the test case source and header.
test.golden_data none Path to an .npz holding the stimulus and the expected outputs.
test.tolerance 1.0 Absolute per-element difference allowed in the output dtype’s units; see tolerance.
test.num_iterations 1 How many times the stimulus is run before the outputs are checked.
test.state_feedback empty (output_index, input_index) pairs carried between iterations.
test.skip_verification false Runs the model but does not compare the outputs.

Configuration is the generated list of every field and its default.

test.golden_data points at a NumPy .npz archive. It is loaded with pickling disabled, so it must hold plain numeric arrays and nothing else.

Key Holds Required
input_0 The stimulus for the model’s first input One key per model input
input_1, input_2, … The stimulus for each further input, in model input order Yes
output_0 The expected value of the model’s first output One key per model output
output_1, output_2, … The expected value of each further output, in model output order Yes

The index in the key is the position of the tensor in the model’s input or output list, not the tensor’s own identifier. Each array is written into that tensor, so its shape and element type are the tensor’s shape and element type.

Provide arrays with exactly the model tensor’s shape and dtype. The loader checks required keys, but do not rely on it to diagnose every shape or dtype mistake in an archive you supply.

A key the conversion expects and does not find fails the conversion with a GoldenDataKeyError naming the key and the file. There is no fallback to another source of truth: the archive supplies the expected values even when skip_verification disables runtime comparison.

For the complete first-deployment path, use the supplied, hash-checked fixture in First conversion. For your own model, first obtain stimulus and expected NumPy arrays from an evaluation you trust, already in the model’s input/output shapes and dtypes. Then this NPZ-writing excerpt saves a single-input, single-output case:

import numpy as np
np.savez(
"golden.npz",
input_0=stimulus,
output_0=expected,
)

This excerpt does not generate inputs or run a reference model. Casting a floating array to int8 is not quantization; apply the model’s scale and zero point in your evaluation pipeline and check the stored values before saving.

Expected data can come from your archive, a host interpreter, or a placeholder when an interpreter is unavailable. Separately, skip_verification decides whether the generated test compares any outputs. These combinations make different claims:

golden_data skip_verification What the expected outputs are What a pass means
Set false The output_N arrays from your file The module matched your expected outputs within tolerance
Set true The file is loaded, but outputs are not compared The supplied stimulus ran without a reported runtime error
Not set false A host interpreter run over deterministic stimulus derived from the model name and tensor The module matched that reference execution within tolerance
Not set true Interpreter outputs if available; otherwise zero placeholders; neither is compared The stimulus ran without a reported runtime error

The no-file, verification-enabled case is a self-consistency check between the generated C and a host interpreter, not a check against data you picked. It catches a lowering that changed behaviour. It cannot catch a model that was wrong before heliaAOT saw it.

Both rows with skip_verification: true are smoke tests. The converter still attempts to create a host interpreter when no golden file is supplied, and may run it if creation succeeds. The flag permits conversion without one; it does not mean that expected outputs will be checked later on the device. The printed Test passed! marker also appears after a successful smoke run; require verification-enabled configuration and per-output comparison records before reporting a numerical match.

The comparison is absolute and per element, in the stored output values. For integer output it compares integer differences and casts test.tolerance to an integer: 1.5 therefore permits a difference of 1. It does not compare dequantized values or relative error. For floating output it compares the absolute difference using float arithmetic and a float tolerance.

The default 1.0 is not an accuracy recommendation for every model. Choose a non-negative, finite tolerance from your numerical requirements and trusted reference, then investigate any mismatch before changing it. A tolerance suitable for quantized integer values may be far too loose for floating outputs.

For floating outputs, an expected NaN matches an actual NaN, and an expected infinity matches only an infinity of the same sign. When the expected value is finite, an actual NaN or infinity fails. This policy checks agreement with the oracle; an oracle containing NaNs still needs review for your application.

num_iterations runs the stimulus through the model repeatedly before the outputs are checked once, at the end. On a stateless model this only exercises the run path; the extra passes change nothing.

Internal persistent state can evolve across invocations without a feedback setting. For a model that exposes state as inputs and outputs, state_feedback names the outputs to copy back into inputs between iterations. Each pair is (output_index, input_index) in model I/O order, must have matching shape and dtype, and may drive an input only once. For example, state_feedback: [[1, 1]] feeds the second output into the second input when those tensors form a compatible state pair.

The host interpreter carries that feedback through the same number of iterations. When you supply an archive instead, its output_N arrays must already contain the expected final-iteration outputs for that stimulus and initial state. The test checks outputs once, after the last iteration. Operators covers the pairs and how the generated code seeds each kind of input.

To repeat the test from fresh persistent state, initialize it again before calling test_case_run. Within one run, num_iterations intentionally carries state without reinitializing the model.

With memory.allocate_arenas: false, your application must load and bind cold constant sidecars. The generated test harness does not perform that load. It returns -1 for external-arena configurations with non-writable cold constant arenas. For cold constants in writable memory it can allocate a buffer, but that buffer does not contain the sidecar weights and cannot establish correct inference. Supply a harness that loads the emitted __blob.bin bytes, binds every region, initializes the model and checks outputs against the same oracle. See Caller-supplied arenas.

int32_t aot_test_case_init(void);
int32_t aot_test_case_run(void);

Both return int32_t and both return 0 on success, so a harness can treat either non-zero return as a failure. They fail for different reasons, and the difference is worth keeping in your logs: a non-zero return from aot_test_case_init is an initialization failure and nothing has run yet, while a non-zero return from aot_test_case_run means either an operator returned an error status or at least one output element exceeded the tolerance.

Check the conversion exit code before building or publishing its output:

Code Meaning
0 The module was written.
1 A typed conversion failure, or an unhandled exception; inspect the error output.
2 The arguments did not validate, with one message per offending field.
130 Interrupted.

Typed conversion failures have a message and may include a remediation hint. Some input failures still raise standard exceptions, including a missing golden archive or unknown transform name. Preserve the traceback and check Troubleshooting before reporting a defect.

Converting into a directory that already holds that module is refused rather than silently overwritten, so a pipeline either cleans its output directory or passes --force deliberately.

Set memory.dump_residency_json and the conversion writes <prefix>_residency.json next to the module. That file, not the console output, is what a pipeline should read.

Terminal window
helia-aot convert --path kws.yaml --memory.dump-residency-json

It carries a schema version, the arena envelope for each role, every tensor with its role, memory, offset and size, and two hashes: plan_hash over the arena envelope, and tensor_layout_hash over per-tensor placement. Gate on it the way you would gate on a linker map: fail the build when an arena grows past what the board has, and treat a hash change as a change that needs reviewing. Memory reads a full report field by field.

Compare the generated source as well as the residency reports. For example, keep baseline and candidate conversions in separate directories:

Terminal window
diff -ru out-baseline/my_model out-candidate/my_model

A difference exits with status 1; an execution error exits with status 2. Review the changed files alongside model/configuration versions before accepting a candidate. A source diff explains code changes, not their performance impact.

That is also the pattern this repository uses on itself. A set of generated files is committed as baselines, the unit suite regenerates them and compares byte for byte, and a change that alters codegen fails until the baselines are updated deliberately. Applied to your own project, it means committing a converted module for one representative model and letting the diff be the review.

The checks below run on heliaAOT itself. They are not checks on your model, and a green run covers those tested cases; it does not validate your conversion. They are worth knowing because they set the floor of what is known to work.

Job What it does
lint-unit Lint and the unit suite, on Python 3.11 and 3.14
e2e Full conversions built and executed for apollo4p_blue_kbr_evb and apollo510_evb on the Corstone-300 fixed virtual platform
zephyr-build Compiles a generated Zephyr module, as a compile check rather than a run

Your own pipeline is the mirror image of that: convert, build, and run the generated test case on the board or on a model of it, then keep the residency report and the emitted tree as artifacts so the next run has something to be compared against.