# On-device validation

Run the application built on the previous page and assess its output. The
walkthrough uses the supplied KWS fixture, verification enabled and a fixed
tolerance of one stored int8 output step.

## What a successful check establishes

| Result | What it establishes |
| --- | --- |
| It compiles and links | The chosen toolchain, headers and libraries can build that firmware. |
| It runs | Model initialization and execution complete successfully in the identified environment. |
| It matches | Output comparison was enabled and every checked element meets the chosen acceptance criterion. |

A single fixture match is a compatibility check. Evaluate application accuracy
separately on representative data. Record whether execution was on a host,
simulator or physical board; an FVP result does not measure Ambiq silicon.

## Configure the test case

The `kws.yaml` from [First conversion](https://ambiqai.github.io/helia-aot/getting-started/convert/) already
contains this test configuration:

```yaml
test:
  enabled: true
  golden_data: golden.npz
  tolerance: 1.0
  num_iterations: 1
  skip_verification: false
```

The supplied fixture and settings are sufficient for this walkthrough. When
using your own data, check the array keys, shapes and types before conversion.

Golden data, tolerance and repeated runs

`input_N` and `output_N` in the NPZ use zero-based positions in the model's input
and output lists. They are not flatbuffer tensor identifiers. Supply every key
with the corresponding tensor's shape and dtype; missing keys fail conversion.
Check shapes and dtypes yourself rather than relying on missing-key validation
to validate the complete fixture.

With supplied golden data, the arrays are embedded directly in the test. Without
it, the converter generates deterministic input and obtains expected output from
a host reference interpreter. For multiple iterations, supplied expected outputs
must represent the final iteration under the same reset/carry protocol.

| Setting | Meaning |
| --- | --- |
| `test.enabled` | Emit the test source and header; default false. |
| `test.golden_data` | Optional NPZ stimulus/expected arrays. |
| `test.tolerance` | Absolute element difference; default 1.0. Integer comparisons cast it to an integer. |
| `test.num_iterations` | Runs before the final comparison; default 1. |
| `test.state_feedback` | Explicit output-to-input carry pairs; default empty. |
| `test.skip_verification` | Omit output comparison; default false. |

For int8, the tolerance is in stored quantized units, not dequantized real units.
For float output, finite expected values reject non-finite actual values; NaN
and infinity expectations are checked by class/sign. Choose a tolerance from
the model's numerical contract before interpreting a result.

## Call it from your application

The [Zephyr application](https://ambiqai.github.io/helia-aot/getting-started/integrate/#zephyr) already
calls `aot_test_case_init`, stops on failure, calls `aot_test_case_run` and prints
both statuses. Use that firmware for this step.

Connect the Apollo510 EVB through its J-Link USB connection. Install the SEGGER
J-Link software and pylink required by Zephyr's runner. This next command
**programs the connected board**; run it only when the build and attached board
are the ones you intend to test:

```sh
west flash -d build/kws_ref
```

Open the board's serial port in a terminal at **115200 baud, 8 data bits, no
parity, one stop bit**, enable log capture, then reset the board. These connection
and console settings follow the
[Zephyr Apollo510 board instructions](https://docs.zephyrproject.org/latest/boards/ambiq/apollo510_evb/doc/index.html#programming-and-debugging).

If you use another module format, enable its console first. `AOT_PRINTF` maps to
`ns_lp_printf` for neuralSPOT and to `printk` for Zephyr with `CONFIG_PRINTK`.
CMake, NSX and pack output default to silent logging unless the application
supplies a hook to generated source compilation.

## What the two functions do

`aot_test_case_init` initializes the test's model context. A nonzero status ends
the test before inference. `aot_test_case_run` copies the stimulus, executes the
configured iterations and stops if the model returns an error. With verification
enabled it then checks all final outputs and returns nonzero on a mismatch.

The harness also reports timing from an available PMU, DWT or SysTick counter,
or reports that a counter was unavailable. Counter output is separate from the
numerical verdict and is not a substitute for controlled board profiling.

For the KWS application, check the following records together. This is an
illustration of the format, not a captured run:

```text
KWS golden check: starting
KWS init status=0
...
HELIA_E2E_TOLERANCE output=0 format=int max_diff=0 tolerance=1
Test passed!
KWS test status=0
```

`max_diff` is the largest observed difference for that output. Require a record
for every model output, the intended tolerance, successful statuses and a saved
configuration with verification enabled. **`Test passed!` alone is insufficient:**
the harness also prints it after a successful run with
`test.skip_verification: true`, when no output was compared.

For float output, `format=f32bits` represents margins as IEEE-754 bit patterns
so target logging does not round a small difference to zero. The
[testing guide](https://ambiqai.github.io/helia-aot/guide/testing/) describes automation and simulator
checks. Repository FVP execution and Zephyr compile checks are different forms
of evidence; neither is a recorded run of your board application.

## When it does not match

Keep the failing logs and acceptance threshold while investigating.

| Observation | Check next |
| --- | --- |
| Nonzero init status | Arena binding/alignment, constant loading and operator initialization; no inference result is available yet. |
| `Model run failed with status N` | The operator/runtime failure before comparison. |
| Missing comparison records | Verification setting, expected output count and logging hooks. |
| Many wrong elements | Model/golden hashes, input order, dtype/shape and preprocessing, then generated and linked module identity. |
| A few small differences | Identify the responsible numeric path and reproduce against the reference; size alone does not establish benign rounding. |
| Host or FVP matches but board differs | Firmware identity, library/toolchain flags, memory placement and target setup, as well as possible implementation defects. |
| Recurrent output differs | Initial state, reset boundaries and explicit state feedback, when the model exposes state I/O. |

Do not raise the tolerance simply to exceed `max_diff`. A justified change to
acceptance criteria needs a numerical explanation and application-level
validation; an unchanged argmax alone is not a general correctness test.

Internal LSTM/SVDF state persists across runs; explicit GRU state is carried
through model I/O. Follow [Stateful models](https://ambiqai.github.io/helia-aot/guide/operators/#stateful-models)
when increasing iterations. The default test is not a sidecar loader for all
external-arena configurations: non-writable cold constant arenas disable its
built-in entry points. Keep the walkthrough's internally allocated arenas for
this first check.

Retain the model and golden hashes, effective configuration, compiler version,
kernel and SDK/toolchain revisions, firmware artifact and raw console log.
Identify the board and execution environment. This lets a later regeneration be
compared against the same evidence.
