Benchmarks and measurement scope
Use this page to assess the scope of recorded results before applying them to your own model. The host-linked KWS record below describes deployment footprints without target validation. A separate KWS comparison records complete profiling-image sizes; the older suite records historical section sizes and inference times. They use different releases and cannot be treated as one controlled experiment. For a current decision, measure your own conversion with its model, target and build settings recorded.
MLPerf Tiny is MLCommons’ benchmark suite for machine learning on deeply embedded devices. It covers five models that between them exercise the operators most small graphs are built from: Add, Average Pooling, Convolution, Depthwise Convolution, Fully Connected, Reshape and Softmax.
- Anomaly Detection (AD)
- Keyword Spotting (KWS)
- Image Classification (IC)
- Streaming Keyword Spotting (STRM)
- Visual Wake Words (VWW)
KWS host-linked deployment
Section titled “KWS host-linked deployment”KWS · AOT 0.20 · host-linked deployment. On 2026-09-25, three host links used the same Apollo510 platform shell, platform archive set and linker script: an empty platform control, a model-specific TFLM / Arm deployment and a historical AOT 0.20 / ns deployment. Profiling and RTT capture storage were absent. These are static build results; target correctness, startup/arena fit, latency, peak live RAM and energy remain untested.
| Metric (bytes) | Empty control | TFLM / Arm | AOT 0.20 / ns |
|---|---|---|---|
| Linked load image (PT_LOAD file bytes) | 52,004 | 245,668 | 100,308 |
| Allocated resident RAM, including ITCM and reserved stack | 19,040 | 53,296 | 40,184 |
| Reserved stack, included above | 16,384 | 16,384 | 16,384 |
The AOT deployment has 59.17% less linked load image and 24.60% less allocated resident RAM in this pair. These are whole-deployment comparisons, not isolated runtime savings: model representation, provider kernels, C/C++ support and runtime differ. An increment over the empty control still describes the incremental deployment, not runtime-only overhead. Heap address reservation is excluded from resident RAM; the fixed stack reservation is not a measured stack watermark.
| Condition | Recorded identity |
|---|---|
| Model | INT8 KWS; input 1×49×10×1, output 1×12 |
| Model SHA-256 | aeea436800704fce17b17292e4412630ad856e9d777c044c64ef748a880bd0ae |
| Compiler | ATfE 22.1.0; M55 hard-float / fast-math; shell -O3, retained interpreter -Os, generated model -O3 |
| Platform | NSX 0.8.1; common Apollo510 platform archives and linker script |
| TFLM fork | 7c1b162c0fd2336876b69daaa20c87a1e7e2f508 |
| Arm provider | 6d21a6f821fb72541173a6c4d05d83329fa74f7c |
| ns provider | aaeb145a67c3decd9869f96474e36e7dbdc2030c |
| TFLM ELF SHA-256 | d1831fc7264ccccf9d45607d3f44ba5848a098e040271130e813a1e7f01a3e25 |
| AOT ELF SHA-256 | 5ad662874d9b76a616af9096872d8aac1bcaed71afc13f0caa9cdebaa0ccaa87 |
Generated model code, arenas and provider/runtime archives were reused from retained builds; the deployment shells and links were new. This is not a clean source rebuild of every dependency. TFLM registers the six operators this model requires. The comparison does not isolate instrumentation removal, and timings from older profiling images cannot be attached to these new binaries.
The extracted figures and source-record hashes are preserved in the deployment data. This download contains the extracted totals and provenance hashes, not the full ELF/map accounting. The underlying KWS deployment record dated 2026-09-25 (comparison.json, ELF/map accounting and RESULTS.md) is retained separately. This record uses a different model identity and toolchain from the profiling comparison below; do not combine the numbers. It does not establish performance for later compiler releases.
Next benchmark step: validate outputs on target for these exact binaries before collecting timing or energy. Peak-memory evidence additionally needs stack watermark and dynamic-lifetime accounting. No target result is implied by a successful host link.
KWS profiling-image comparison
Section titled “KWS profiling-image comparison”A hardware-validation run on 2026-09-16 built the same KWS model with heliaRT 1.20.0 and heliaAOT 0.20.0. GNU size reported 278,596 → 151,160 bytes of code + read-only data, a 45.7% reduction for these profiling images.
The metric is GNU size’s text aggregate, which includes code and read-only data. It describes the complete profiling firmware, including its harness. It is not the literal ELF .text section, model-only code size, total flash usage, or peak runtime memory.
| Condition | Both builds |
|---|---|
| Model | kws_ref_model.tflite, 53,744 bytes |
| Model SHA-256 | 935c513a5f4dbfe6458c16b0a791e916f5e89d6c9b02ffbefce49c06fbbdf038 |
| Target | Apollo510 EVB, Cortex-M55, LP 96 MHz |
| Compiler | Arm GNU Toolchain 15.2.Rel1, GCC 15.2.1; CMake 4.1.2 |
| Runtime variant | heliaRT uses release-with-logs; AOT emits a static module |
| Profiling | hpx 0.1.6; RTT; CPU counters; 3 iterations, 1 warmup |
| Placement | Default linker profile; no arena or weights-location override; actual region usage differs by engine |
| Kernels | ns-cmsis-nn aaeb145a67c3decd9869f96474e36e7dbdc2030c |
| GNU size metric (bytes) | heliaRT 1.20.0 | heliaAOT 0.20.0 |
|---|---|---|
| Text: code + read-only data | 278,596 | 151,160 |
| Data | 154,796 | 2,752 |
| BSS | 86,260 | 185,412 |
Both runs enable per-layer profiling. The runtime uses release-with-logs; the generated AOT module has no matching runtime-variant setting, so logging and build modes are not held identical. This comparison does not isolate compilation as the sole cause of the difference.
The smaller text aggregate does not mean every memory category shrinks: BSS grows in this pair. These results describe the recorded builds, rather than predicting the size of another model or application. No latency or energy improvement is claimed here.
The hardware-validation run and recorded result bundles contain each case’s summary.json, run_metadata.json and nsx.lock. The matching cases are apollo510_evb-kws-rt-ns-arm-none-eabi-gcc-rtt-auto and apollo510_evb-kws-aot-ns-arm-none-eabi-gcc-rtt-auto. The local benchmark data preserves the extracted figures and source-file hashes.
Historical MLPerf Tiny suite
Section titled “Historical MLPerf Tiny suite”The remaining results reproduce an older report last updated on 2025-08-19. That date is the report’s update date; the original source does not establish a measurement date. Its model hashes, precision, compiler options, clocks, placement and raw build receipts were not preserved alongside the table. Treat it as a historical record, not a controlled comparison of current releases.
What was measured
Section titled “What was measured”Each model was built twice for the same board: once with the interpreter-based engine, which carries the runtime and its operator resolver, and once with heliaAOT, which carries the kernels the model uses and its planned arenas. Text, data and BSS are the sections of the built image; the inference time is for one inference.
Measured. heliaRT v1.3.0 against heliaAOT v0.2.2 on the Apollo510 EVB; historical report updated 2025-08-19. Reductions below are computed from the reported figures.
The original report describes matching models and board. Missing build controls limit which differences can be attributed to the engine.
Results
Section titled “Results”| Model | Build | Text (KB) | Data (KB) | BSS (KB) | Total (KB) | Inference (µs) |
|---|---|---|---|---|---|---|
| AD | heliaRT | 201 | 279 | 117 | 596 | 292 |
| AD | heliaAOT | 58 | 272 | 30 | 360 | 275 |
| KWS | heliaRT | 201 | 61 | 137 | 399 | 8,085 |
| KWS | heliaAOT | 83 | 31 | 59 | 173 | 8,074 |
| IC | heliaRT | 211 | 105 | 163 | 478 | 20,139 |
| IC | heliaAOT | 78 | 84 | 79 | 241 | 20,041 |
| STRM | heliaRT | 211 | 82 | 208 | 500 | 1,730 |
| STRM | heliaAOT | 84 | 54 | 52 | 189 | 1,719 |
| VWW | heliaRT | 200 | 334 | 217 | 751 | 29,480 |
| VWW | heliaAOT | 107 | 222 | 163 | 491 | 29,327 |
Reported binary sections by model and build
Text + data + BSS, in reported KB
- heliaRT v1.3.0
- heliaAOT v0.2.2
What that comes to
Section titled “What that comes to”The historical table estimates flash as text plus data and labels BSS as RAM; BSS is not peak runtime RAM usage. The total is the reported section aggregate. The throughput ratio is the heliaRT inference time divided by the heliaAOT one, so above 1.00 is faster.
| Model | Flash | RAM | Total | Throughput ratio |
|---|---|---|---|---|
| AD | 480 to 330 KB, 31% smaller | 117 to 30 KB, 74% smaller | 596 to 360 KB, 40% smaller | 1.06× |
| KWS | 262 to 114 KB, 56% smaller | 137 to 59 KB, 57% smaller | 399 to 173 KB, 57% smaller | 1.00× |
| IC | 316 to 162 KB, 49% smaller | 163 to 79 KB, 52% smaller | 478 to 241 KB, 50% smaller | 1.00× |
| STRM | 293 to 138 KB, 53% smaller | 208 to 52 KB, 75% smaller | 500 to 189 KB, 62% smaller | 1.01× |
| VWW | 534 to 329 KB, 38% smaller | 217 to 163 KB, 25% smaller | 751 to 491 KB, 35% smaller | 1.01× |
Across the suite. Flash falls 31% to 56%, median 49%. RAM falls 25% to 75%, median 57%. Total footprint falls 35% to 62%, median 50%. Throughput moves between 1.00× and 1.06×.
Both tables and the chart are written from
src/data/benchmarks/mlperf-tiny.json, which holds the measured figures and
nothing derived: the reductions above are computed from them on every build.
Measure your own
Section titled “Measure your own”The comparison above is one engine against another on one board. To run the same kind of comparison on your model, see Measure a conversion.