# Testing and automation

Use the generated test harness to compare a module with trusted outputs, then
repeat conversion, build and validation in your pipeline. This guide covers
golden fixtures, tolerance and state handling, followed by exit codes and
reports for automation.

[Validate it on target](https://ambiqai.github.io/helia-aot/getting-started/validate/) is the
walkthrough for a first module. This page is the detail behind it.

## Start with one checked inference

Use a model that already converts and a target build that can call its generated
C API. Add `test.enabled: true` to the conversion configuration. If you have
trusted expected outputs, supply a golden archive; otherwise the converter
needs a host interpreter for the model.

```bash
helia-aot convert --path convert.yaml --test.enabled
```

Build the generated test source with the module in your target project, then
call `aot_test_case_init()` once and `aot_test_case_run()`. With the default
prefix `aot`, both functions return `0` on success. Check both results; an
initialization failure must stop the harness. The [first validation
walkthrough](https://ambiqai.github.io/helia-aot/getting-started/validate/) supplies the integration path.

With output verification enabled, a successful run means the supplied stimulus
produced outputs within the configured tolerance of the chosen oracle. It does
not establish model accuracy over a dataset or correctness for every input.
Keep the oracle, tolerance, model, configuration and target build identity with
the result before changing conversion controls. This shortest path uses the
default internally allocated arenas; external buffers have the additional
[harness requirements](https://ambiqai.github.io/helia-aot/guide/testing/#external-arena-tests) below.

## The generated test case

Set `test.enabled` and the conversion emits `aot_test_case.c` and its header
alongside the module, with the stimulus and the expected outputs baked in as
constant arrays. The test needs no file system on the board and reads nothing
at run time.

```yaml title="kws.yaml"
test:
  enabled: true
  golden_data: golden.npz
  tolerance: 1.0
  num_iterations: 1
  skip_verification: false
```

| Field | Default | What it does |
| --- | --- | --- |
| `test.enabled` | `false` | Emits the test case source and header. |
| `test.golden_data` | none | Path to an `.npz` holding the stimulus and the expected outputs. |
| `test.tolerance` | `1.0` | Absolute per-element difference allowed in the output dtype's units; see [tolerance](https://ambiqai.github.io/helia-aot/guide/testing/#choosing-a-tolerance). |
| `test.num_iterations` | `1` | How many times the stimulus is run before the outputs are checked. |
| `test.state_feedback` | empty | `(output_index, input_index)` pairs carried between iterations. |
| `test.skip_verification` | `false` | Runs the model but does not compare the outputs. |

[Configuration](https://ambiqai.github.io/helia-aot/reference/configuration/) is the generated list of
every field and its default.

### The golden file

`test.golden_data` points at a NumPy `.npz` archive. It is loaded with pickling
disabled, so it must hold plain numeric arrays and nothing else.

| Key | Holds | Required |
| --- | --- | --- |
| `input_0` | The stimulus for the model's first input | One key per model input |
| `input_1`, `input_2`, ... | The stimulus for each further input, in model input order | Yes |
| `output_0` | The expected value of the model's first output | One key per model output |
| `output_1`, `output_2`, ... | The expected value of each further output, in model output order | Yes |

The index in the key is the position of the tensor in the model's input or
output list, not the tensor's own identifier. Each array is written into that
tensor, so its shape and element type are the tensor's shape and element type.

Provide arrays with exactly the model tensor's shape and dtype. The loader
checks required keys, but do not rely on it to diagnose every shape or dtype
mistake in an archive you supply.

A key the conversion expects and does not find fails the conversion with a
`GoldenDataKeyError` naming the key and the file. There is no fallback to
another source of truth: the archive supplies the expected values even when
`skip_verification` disables runtime comparison.

For the complete first-deployment path, use the supplied, hash-checked fixture
in [First conversion](https://ambiqai.github.io/helia-aot/getting-started/convert/#get-a-model).
For your own model, first obtain `stimulus` and `expected` NumPy arrays from
an evaluation you trust, already in the model's input/output shapes and dtypes.
Then this **NPZ-writing excerpt** saves a single-input, single-output case:

```python
import numpy as np

np.savez(
    "golden.npz",
    input_0=stimulus,
    output_0=expected,
)
```

This excerpt does not generate inputs or run a reference model. Casting a
floating array to int8 is not quantization; apply the model's scale and zero
point in your evaluation pipeline and check the stored values before saving.

### The three oracles

Expected data can come from your archive, a host interpreter, or a placeholder
when an interpreter is unavailable. Separately, `skip_verification` decides
whether the generated test compares any outputs. These combinations make
different claims:

| `golden_data` | `skip_verification` | What the expected outputs are | What a pass means |
| --- | --- | --- | --- |
| Set | `false` | The `output_N` arrays from your file | The module matched your expected outputs within tolerance |
| Set | `true` | The file is loaded, but outputs are not compared | The supplied stimulus ran without a reported runtime error |
| Not set | `false` | A host interpreter run over deterministic stimulus derived from the model name and tensor | The module matched that reference execution within tolerance |
| Not set | `true` | Interpreter outputs if available; otherwise zero placeholders; neither is compared | The stimulus ran without a reported runtime error |

The no-file, verification-enabled case is a
self-consistency check between the generated C and a host interpreter, not a
check against data you picked. It catches a lowering that changed behaviour. It
cannot catch a model that was wrong before heliaAOT saw it.

Both rows with `skip_verification: true` are smoke tests. The converter still
attempts to create a host interpreter when no golden file is supplied, and may
run it if creation succeeds. The flag permits conversion without one; it does
not mean that expected outputs will be checked later on the device. The
printed `Test passed!` marker also appears after a successful smoke run; require
verification-enabled configuration and per-output comparison records before
reporting a numerical match.

:::caution
The interpreter oracle compares the module against a host interpreter run of the
same model, so it catches lowering and placement mistakes but it cannot catch a
model that was quantized badly in the first place. Both sides would be wrong in
the same way. For anything you are shipping, supply `golden_data` produced by
the evaluation you already trust.
:::

### Choosing a tolerance

The comparison is absolute and per element, in the stored output values.
For integer output it compares integer differences and casts `test.tolerance`
to an integer: `1.5` therefore permits a difference of `1`. It does not compare
dequantized values or relative error. For floating output it compares the
absolute difference using float arithmetic and a float tolerance.

The default `1.0` is not an accuracy recommendation for every model. Choose a
non-negative, finite tolerance from your numerical requirements and trusted
reference, then investigate any mismatch before changing it. A tolerance
suitable for quantized integer values may be far too loose for floating outputs.

For floating outputs, an expected NaN matches an actual NaN, and an expected
infinity matches only an infinity of the same sign. When the expected value is
finite, an actual NaN or infinity fails. This policy checks agreement with the
oracle; an oracle containing NaNs still needs review for your application.

### Iterations and state

`num_iterations` runs the stimulus through the model repeatedly before the
outputs are checked once, at the end. On a stateless model this only exercises
the run path; the extra passes change nothing.

Internal persistent state can evolve across invocations without a feedback
setting. For a model that exposes state as inputs and outputs, `state_feedback`
names the outputs to copy back into inputs between iterations. Each pair is
`(output_index, input_index)` in model I/O order, must have matching shape and
dtype, and may drive an input only once. For example, `state_feedback: [[1, 1]]`
feeds the second output into the second input when those tensors form a
compatible state pair.

The host interpreter carries that feedback through the same number of
iterations. When you supply an archive instead, its `output_N` arrays must
already contain the expected **final-iteration** outputs for that stimulus
and initial state. The test checks outputs once, after the last iteration.
[Operators](https://ambiqai.github.io/helia-aot/guide/operators/#carrying-state-through-a-test-run)
covers the pairs and how the generated code seeds each kind of input.

To repeat the test from fresh persistent state, initialize it again before
calling `test_case_run`. Within one run, `num_iterations` intentionally carries
state without reinitializing the model.

### External-arena tests

With `memory.allocate_arenas: false`, your application must load and bind cold
constant sidecars. The generated test harness does not perform that load. It
returns `-1` for external-arena configurations with non-writable cold constant
arenas. For cold constants in writable memory it can allocate a buffer, but
that buffer does not contain the sidecar weights and cannot establish correct
inference. Supply a harness that loads the emitted `__blob.bin` bytes, binds
every region, initializes the model and checks outputs against the same oracle.
See [Caller-supplied arenas](https://ambiqai.github.io/helia-aot/guide/memory-placement/#caller-supplied-arenas).

### The two functions

```c
int32_t aot_test_case_init(void);
int32_t aot_test_case_run(void);
```

Both return `int32_t` and both return 0 on success, so a harness can treat
either non-zero return as a failure. They fail for different reasons, and the
difference is worth keeping in your logs: a non-zero return from
`aot_test_case_init` is an initialization failure and nothing has run yet,
while a non-zero return from `aot_test_case_run` means either an operator
returned an error status or at least one output element exceeded the tolerance.

## Automating conversions

### Exit codes

Check the conversion exit code before building or publishing its output:

| Code | Meaning |
| --- | --- |
| 0 | The module was written. |
| 1 | A typed conversion failure, or an unhandled exception; inspect the error output. |
| 2 | The arguments did not validate, with one message per offending field. |
| 130 | Interrupted. |

Typed conversion failures have a message and may include a remediation hint. Some input failures still
raise standard exceptions, including a missing golden archive or unknown
transform name. Preserve the traceback and check
[Troubleshooting](https://ambiqai.github.io/helia-aot/guide/troubleshooting/) before reporting a defect.

Converting into a directory that already holds that module is refused rather
than silently overwritten, so a pipeline either cleans its output directory or
passes `--force` deliberately.

### The machine-readable report

Set `memory.dump_residency_json` and the conversion writes
`<prefix>_residency.json` next to the module. That file, not the console
output, is what a pipeline should read.

```sh
helia-aot convert --path kws.yaml --memory.dump-residency-json
```

It carries a schema version, the arena envelope for each role, every tensor
with its role, memory, offset and size, and two hashes: `plan_hash` over the
arena envelope, and `tensor_layout_hash` over per-tensor placement. Gate on it
the way you would gate on a linker map: fail the build when an arena grows past
what the board has, and treat a hash change as a change that needs reviewing.
[Memory](https://ambiqai.github.io/helia-aot/guide/memory/#the-residency-report) reads a full report
field by field.

### Comparing two conversions

Compare the generated source as well as the residency reports. For example,
keep baseline and candidate conversions in separate directories:

```sh
diff -ru out-baseline/my_model out-candidate/my_model
```

A difference exits with status 1; an execution error exits with status 2. Review
the changed files alongside model/configuration versions before accepting a
candidate. A source diff explains code changes, not their performance impact.

That is also the pattern this repository uses on itself. A set of generated
files is committed as baselines, the unit suite regenerates them and compares
byte for byte, and a change that alters codegen fails until the baselines are
updated deliberately. Applied to your own project, it means committing a
converted module for one representative model and letting the diff be the
review.

## What this repository's CI covers

The checks below run on heliaAOT itself. They are not checks on your model, and
a green run covers those tested cases; it does not validate your conversion.
They are worth knowing because they set the floor of what is known to work.

| Job | What it does |
| --- | --- |
| `lint-unit` | Lint and the unit suite, on Python 3.11 and 3.14 |
| `e2e` | Full conversions built and executed for `apollo4p_blue_kbr_evb` and `apollo510_evb` on the Corstone-300 fixed virtual platform |
| `zephyr-build` | Compiles a generated Zephyr module, as a compile check rather than a run |

Your own pipeline is the mirror image of that: convert, build, and run the
generated test case on the board or on a model of it, then keep the residency
report and the emitted tree as artifacts so the next run has something to be
compared against.
