Skip to content
heliaAOT
HELIA HUB

On-device validation

Run the application built on the previous page and assess its output. The walkthrough uses the supplied KWS fixture, verification enabled and a fixed tolerance of one stored int8 output step.

Result What it establishes
It compiles and links The chosen toolchain, headers and libraries can build that firmware.
It runs Model initialization and execution complete successfully in the identified environment.
It matches Output comparison was enabled and every checked element meets the chosen acceptance criterion.

A single fixture match is a compatibility check. Evaluate application accuracy separately on representative data. Record whether execution was on a host, simulator or physical board; an FVP result does not measure Ambiq silicon.

The kws.yaml from First conversion already contains this test configuration:

test:
enabled: true
golden_data: golden.npz
tolerance: 1.0
num_iterations: 1
skip_verification: false

The supplied fixture and settings are sufficient for this walkthrough. When using your own data, check the array keys, shapes and types before conversion.

Golden data, tolerance and repeated runs

input_N and output_N in the NPZ use zero-based positions in the model’s input and output lists. They are not flatbuffer tensor identifiers. Supply every key with the corresponding tensor’s shape and dtype; missing keys fail conversion. Check shapes and dtypes yourself rather than relying on missing-key validation to validate the complete fixture.

With supplied golden data, the arrays are embedded directly in the test. Without it, the converter generates deterministic input and obtains expected output from a host reference interpreter. For multiple iterations, supplied expected outputs must represent the final iteration under the same reset/carry protocol.

Setting Meaning
test.enabled Emit the test source and header; default false.
test.golden_data Optional NPZ stimulus/expected arrays.
test.tolerance Absolute element difference; default 1.0. Integer comparisons cast it to an integer.
test.num_iterations Runs before the final comparison; default 1.
test.state_feedback Explicit output-to-input carry pairs; default empty.
test.skip_verification Omit output comparison; default false.

For int8, the tolerance is in stored quantized units, not dequantized real units. For float output, finite expected values reject non-finite actual values; NaN and infinity expectations are checked by class/sign. Choose a tolerance from the model’s numerical contract before interpreting a result.

The Zephyr application already calls aot_test_case_init, stops on failure, calls aot_test_case_run and prints both statuses. Use that firmware for this step.

Connect the Apollo510 EVB through its J-Link USB connection. Install the SEGGER J-Link software and pylink required by Zephyr’s runner. This next command programs the connected board; run it only when the build and attached board are the ones you intend to test:

Terminal window
west flash -d build/kws_ref

Open the board’s serial port in a terminal at 115200 baud, 8 data bits, no parity, one stop bit, enable log capture, then reset the board. These connection and console settings follow the Zephyr Apollo510 board instructions.

If you use another module format, enable its console first. AOT_PRINTF maps to ns_lp_printf for neuralSPOT and to printk for Zephyr with CONFIG_PRINTK. CMake, NSX and pack output default to silent logging unless the application supplies a hook to generated source compilation.

aot_test_case_init initializes the test’s model context. A nonzero status ends the test before inference. aot_test_case_run copies the stimulus, executes the configured iterations and stops if the model returns an error. With verification enabled it then checks all final outputs and returns nonzero on a mismatch.

The harness also reports timing from an available PMU, DWT or SysTick counter, or reports that a counter was unavailable. Counter output is separate from the numerical verdict and is not a substitute for controlled board profiling.

For the KWS application, check the following records together. This is an illustration of the format, not a captured run:

KWS golden check: starting
KWS init status=0
...
HELIA_E2E_TOLERANCE output=0 format=int max_diff=0 tolerance=1
Test passed!
KWS test status=0

max_diff is the largest observed difference for that output. Require a record for every model output, the intended tolerance, successful statuses and a saved configuration with verification enabled. Test passed! alone is insufficient: the harness also prints it after a successful run with test.skip_verification: true, when no output was compared.

For float output, format=f32bits represents margins as IEEE-754 bit patterns so target logging does not round a small difference to zero. The testing guide describes automation and simulator checks. Repository FVP execution and Zephyr compile checks are different forms of evidence; neither is a recorded run of your board application.

Keep the failing logs and acceptance threshold while investigating.

Observation Check next
Nonzero init status Arena binding/alignment, constant loading and operator initialization; no inference result is available yet.
Model run failed with status N The operator/runtime failure before comparison.
Missing comparison records Verification setting, expected output count and logging hooks.
Many wrong elements Model/golden hashes, input order, dtype/shape and preprocessing, then generated and linked module identity.
A few small differences Identify the responsible numeric path and reproduce against the reference; size alone does not establish benign rounding.
Host or FVP matches but board differs Firmware identity, library/toolchain flags, memory placement and target setup, as well as possible implementation defects.
Recurrent output differs Initial state, reset boundaries and explicit state feedback, when the model exposes state I/O.

Do not raise the tolerance simply to exceed max_diff. A justified change to acceptance criteria needs a numerical explanation and application-level validation; an unchanged argmax alone is not a general correctness test.

Internal LSTM/SVDF state persists across runs; explicit GRU state is carried through model I/O. Follow Stateful models when increasing iterations. The default test is not a sidecar loader for all external-arena configurations: non-writable cold constant arenas disable its built-in entry points. Keep the walkthrough’s internally allocated arenas for this first check.

Retain the model and golden hashes, effective configuration, compiler version, kernel and SDK/toolchain revisions, firmware artifact and raw console log. Identify the board and execution environment. This lets a later regeneration be compared against the same evidence.