On-device validation
Run the application built on the previous page and assess its output. The walkthrough uses the supplied KWS fixture, verification enabled and a fixed tolerance of one stored int8 output step.
What a successful check establishes
Section titled “What a successful check establishes”| Result | What it establishes |
|---|---|
| It compiles and links | The chosen toolchain, headers and libraries can build that firmware. |
| It runs | Model initialization and execution complete successfully in the identified environment. |
| It matches | Output comparison was enabled and every checked element meets the chosen acceptance criterion. |
A single fixture match is a compatibility check. Evaluate application accuracy separately on representative data. Record whether execution was on a host, simulator or physical board; an FVP result does not measure Ambiq silicon.
Configure the test case
Section titled “Configure the test case”The kws.yaml from First conversion already
contains this test configuration:
test: enabled: true golden_data: golden.npz tolerance: 1.0 num_iterations: 1 skip_verification: falseThe supplied fixture and settings are sufficient for this walkthrough. When using your own data, check the array keys, shapes and types before conversion.
Golden data, tolerance and repeated runs
input_N and output_N in the NPZ use zero-based positions in the model’s input
and output lists. They are not flatbuffer tensor identifiers. Supply every key
with the corresponding tensor’s shape and dtype; missing keys fail conversion.
Check shapes and dtypes yourself rather than relying on missing-key validation
to validate the complete fixture.
With supplied golden data, the arrays are embedded directly in the test. Without it, the converter generates deterministic input and obtains expected output from a host reference interpreter. For multiple iterations, supplied expected outputs must represent the final iteration under the same reset/carry protocol.
| Setting | Meaning |
|---|---|
test.enabled |
Emit the test source and header; default false. |
test.golden_data |
Optional NPZ stimulus/expected arrays. |
test.tolerance |
Absolute element difference; default 1.0. Integer comparisons cast it to an integer. |
test.num_iterations |
Runs before the final comparison; default 1. |
test.state_feedback |
Explicit output-to-input carry pairs; default empty. |
test.skip_verification |
Omit output comparison; default false. |
For int8, the tolerance is in stored quantized units, not dequantized real units. For float output, finite expected values reject non-finite actual values; NaN and infinity expectations are checked by class/sign. Choose a tolerance from the model’s numerical contract before interpreting a result.
Call it from your application
Section titled “Call it from your application”The Zephyr application already
calls aot_test_case_init, stops on failure, calls aot_test_case_run and prints
both statuses. Use that firmware for this step.
Connect the Apollo510 EVB through its J-Link USB connection. Install the SEGGER J-Link software and pylink required by Zephyr’s runner. This next command programs the connected board; run it only when the build and attached board are the ones you intend to test:
west flash -d build/kws_refOpen the board’s serial port in a terminal at 115200 baud, 8 data bits, no parity, one stop bit, enable log capture, then reset the board. These connection and console settings follow the Zephyr Apollo510 board instructions.
If you use another module format, enable its console first. AOT_PRINTF maps to
ns_lp_printf for neuralSPOT and to printk for Zephyr with CONFIG_PRINTK.
CMake, NSX and pack output default to silent logging unless the application
supplies a hook to generated source compilation.
What the two functions do
Section titled “What the two functions do”aot_test_case_init initializes the test’s model context. A nonzero status ends
the test before inference. aot_test_case_run copies the stimulus, executes the
configured iterations and stops if the model returns an error. With verification
enabled it then checks all final outputs and returns nonzero on a mismatch.
The harness also reports timing from an available PMU, DWT or SysTick counter, or reports that a counter was unavailable. Counter output is separate from the numerical verdict and is not a substitute for controlled board profiling.
For the KWS application, check the following records together. This is an illustration of the format, not a captured run:
KWS golden check: startingKWS init status=0...HELIA_E2E_TOLERANCE output=0 format=int max_diff=0 tolerance=1Test passed!KWS test status=0max_diff is the largest observed difference for that output. Require a record
for every model output, the intended tolerance, successful statuses and a saved
configuration with verification enabled. Test passed! alone is insufficient:
the harness also prints it after a successful run with
test.skip_verification: true, when no output was compared.
For float output, format=f32bits represents margins as IEEE-754 bit patterns
so target logging does not round a small difference to zero. The
testing guide describes automation and simulator
checks. Repository FVP execution and Zephyr compile checks are different forms
of evidence; neither is a recorded run of your board application.
When it does not match
Section titled “When it does not match”Keep the failing logs and acceptance threshold while investigating.
| Observation | Check next |
|---|---|
| Nonzero init status | Arena binding/alignment, constant loading and operator initialization; no inference result is available yet. |
Model run failed with status N |
The operator/runtime failure before comparison. |
| Missing comparison records | Verification setting, expected output count and logging hooks. |
| Many wrong elements | Model/golden hashes, input order, dtype/shape and preprocessing, then generated and linked module identity. |
| A few small differences | Identify the responsible numeric path and reproduce against the reference; size alone does not establish benign rounding. |
| Host or FVP matches but board differs | Firmware identity, library/toolchain flags, memory placement and target setup, as well as possible implementation defects. |
| Recurrent output differs | Initial state, reset boundaries and explicit state feedback, when the model exposes state I/O. |
Do not raise the tolerance simply to exceed max_diff. A justified change to
acceptance criteria needs a numerical explanation and application-level
validation; an unchanged argmax alone is not a general correctness test.
Internal LSTM/SVDF state persists across runs; explicit GRU state is carried through model I/O. Follow Stateful models when increasing iterations. The default test is not a sidecar loader for all external-arena configurations: non-writable cold constant arenas disable its built-in entry points. Keep the walkthrough’s internally allocated arenas for this first check.
Retain the model and golden hashes, effective configuration, compiler version, kernel and SDK/toolchain revisions, firmware artifact and raw console log. Identify the board and execution environment. This lets a later regeneration be compared against the same evidence.