Skip to content
heliaAOT
HELIA HUB

Benchmarks and measurement scope

Use this page to assess the scope of recorded results before applying them to your own model. The host-linked KWS record below describes deployment footprints without target validation. A separate KWS comparison records complete profiling-image sizes; the older suite records historical section sizes and inference times. They use different releases and cannot be treated as one controlled experiment. For a current decision, measure your own conversion with its model, target and build settings recorded.

MLPerf Tiny is MLCommons’ benchmark suite for machine learning on deeply embedded devices. It covers five models that between them exercise the operators most small graphs are built from: Add, Average Pooling, Convolution, Depthwise Convolution, Fully Connected, Reshape and Softmax.

KWS · AOT 0.20 · host-linked deployment. On 2026-09-25, three host links used the same Apollo510 platform shell, platform archive set and linker script: an empty platform control, a model-specific TFLM / Arm deployment and a historical AOT 0.20 / ns deployment. Profiling and RTT capture storage were absent. These are static build results; target correctness, startup/arena fit, latency, peak live RAM and energy remain untested.

Metric (bytes) Empty control TFLM / Arm AOT 0.20 / ns
Linked load image (PT_LOAD file bytes) 52,004 245,668 100,308
Allocated resident RAM, including ITCM and reserved stack 19,040 53,296 40,184
Reserved stack, included above 16,384 16,384 16,384

The AOT deployment has 59.17% less linked load image and 24.60% less allocated resident RAM in this pair. These are whole-deployment comparisons, not isolated runtime savings: model representation, provider kernels, C/C++ support and runtime differ. An increment over the empty control still describes the incremental deployment, not runtime-only overhead. Heap address reservation is excluded from resident RAM; the fixed stack reservation is not a measured stack watermark.

Condition Recorded identity
Model INT8 KWS; input 1×49×10×1, output 1×12
Model SHA-256 aeea436800704fce17b17292e4412630ad856e9d777c044c64ef748a880bd0ae
Compiler ATfE 22.1.0; M55 hard-float / fast-math; shell -O3, retained interpreter -Os, generated model -O3
Platform NSX 0.8.1; common Apollo510 platform archives and linker script
TFLM fork 7c1b162c0fd2336876b69daaa20c87a1e7e2f508
Arm provider 6d21a6f821fb72541173a6c4d05d83329fa74f7c
ns provider aaeb145a67c3decd9869f96474e36e7dbdc2030c
TFLM ELF SHA-256 d1831fc7264ccccf9d45607d3f44ba5848a098e040271130e813a1e7f01a3e25
AOT ELF SHA-256 5ad662874d9b76a616af9096872d8aac1bcaed71afc13f0caa9cdebaa0ccaa87

Generated model code, arenas and provider/runtime archives were reused from retained builds; the deployment shells and links were new. This is not a clean source rebuild of every dependency. TFLM registers the six operators this model requires. The comparison does not isolate instrumentation removal, and timings from older profiling images cannot be attached to these new binaries.

The extracted figures and source-record hashes are preserved in the deployment data. This download contains the extracted totals and provenance hashes, not the full ELF/map accounting. The underlying KWS deployment record dated 2026-09-25 (comparison.json, ELF/map accounting and RESULTS.md) is retained separately. This record uses a different model identity and toolchain from the profiling comparison below; do not combine the numbers. It does not establish performance for later compiler releases.

Next benchmark step: validate outputs on target for these exact binaries before collecting timing or energy. Peak-memory evidence additionally needs stack watermark and dynamic-lifetime accounting. No target result is implied by a successful host link.

A hardware-validation run on 2026-09-16 built the same KWS model with heliaRT 1.20.0 and heliaAOT 0.20.0. GNU size reported 278,596 → 151,160 bytes of code + read-only data, a 45.7% reduction for these profiling images.

The metric is GNU size’s text aggregate, which includes code and read-only data. It describes the complete profiling firmware, including its harness. It is not the literal ELF .text section, model-only code size, total flash usage, or peak runtime memory.

Condition Both builds
Model kws_ref_model.tflite, 53,744 bytes
Model SHA-256 935c513a5f4dbfe6458c16b0a791e916f5e89d6c9b02ffbefce49c06fbbdf038
Target Apollo510 EVB, Cortex-M55, LP 96 MHz
Compiler Arm GNU Toolchain 15.2.Rel1, GCC 15.2.1; CMake 4.1.2
Runtime variant heliaRT uses release-with-logs; AOT emits a static module
Profiling hpx 0.1.6; RTT; CPU counters; 3 iterations, 1 warmup
Placement Default linker profile; no arena or weights-location override; actual region usage differs by engine
Kernels ns-cmsis-nn aaeb145a67c3decd9869f96474e36e7dbdc2030c
GNU size metric (bytes) heliaRT 1.20.0 heliaAOT 0.20.0
Text: code + read-only data 278,596 151,160
Data 154,796 2,752
BSS 86,260 185,412

Both runs enable per-layer profiling. The runtime uses release-with-logs; the generated AOT module has no matching runtime-variant setting, so logging and build modes are not held identical. This comparison does not isolate compilation as the sole cause of the difference.

The smaller text aggregate does not mean every memory category shrinks: BSS grows in this pair. These results describe the recorded builds, rather than predicting the size of another model or application. No latency or energy improvement is claimed here.

The hardware-validation run and recorded result bundles contain each case’s summary.json, run_metadata.json and nsx.lock. The matching cases are apollo510_evb-kws-rt-ns-arm-none-eabi-gcc-rtt-auto and apollo510_evb-kws-aot-ns-arm-none-eabi-gcc-rtt-auto. The local benchmark data preserves the extracted figures and source-file hashes.

The remaining results reproduce an older report last updated on 2025-08-19. That date is the report’s update date; the original source does not establish a measurement date. Its model hashes, precision, compiler options, clocks, placement and raw build receipts were not preserved alongside the table. Treat it as a historical record, not a controlled comparison of current releases.

Each model was built twice for the same board: once with the interpreter-based engine, which carries the runtime and its operator resolver, and once with heliaAOT, which carries the kernels the model uses and its planned arenas. Text, data and BSS are the sections of the built image; the inference time is for one inference.

Measured. heliaRT v1.3.0 against heliaAOT v0.2.2 on the Apollo510 EVB; historical report updated 2025-08-19. Reductions below are computed from the reported figures.

The original report describes matching models and board. Missing build controls limit which differences can be attributed to the engine.

Model Build Text (KB) Data (KB) BSS (KB) Total (KB) Inference (µs)
AD heliaRT 201 279 117 596 292
AD heliaAOT 58 272 30 360 275
KWS heliaRT 201 61 137 399 8,085
KWS heliaAOT 83 31 59 173 8,074
IC heliaRT 211 105 163 478 20,139
IC heliaAOT 78 84 79 241 20,041
STRM heliaRT 211 82 208 500 1,730
STRM heliaAOT 84 54 52 189 1,719
VWW heliaRT 200 334 217 751 29,480
VWW heliaAOT 107 222 163 491 29,327

Reported binary sections by model and build

Text + data + BSS, in reported KB

  • heliaRT v1.3.0
  • heliaAOT v0.2.2
ADKWSICSTRMVWW0200400600800596360399173478241500189751491
MLPerf Tiny on the Apollo510 EVB, historical report updated 2025-08-19. The same figures as the table.

The historical table estimates flash as text plus data and labels BSS as RAM; BSS is not peak runtime RAM usage. The total is the reported section aggregate. The throughput ratio is the heliaRT inference time divided by the heliaAOT one, so above 1.00 is faster.

Model Flash RAM Total Throughput ratio
AD 480 to 330 KB, 31% smaller 117 to 30 KB, 74% smaller 596 to 360 KB, 40% smaller 1.06×
KWS 262 to 114 KB, 56% smaller 137 to 59 KB, 57% smaller 399 to 173 KB, 57% smaller 1.00×
IC 316 to 162 KB, 49% smaller 163 to 79 KB, 52% smaller 478 to 241 KB, 50% smaller 1.00×
STRM 293 to 138 KB, 53% smaller 208 to 52 KB, 75% smaller 500 to 189 KB, 62% smaller 1.01×
VWW 534 to 329 KB, 38% smaller 217 to 163 KB, 25% smaller 751 to 491 KB, 35% smaller 1.01×

Across the suite. Flash falls 31% to 56%, median 49%. RAM falls 25% to 75%, median 57%. Total footprint falls 35% to 62%, median 50%. Throughput moves between 1.00× and 1.06×.

Both tables and the chart are written from src/data/benchmarks/mlperf-tiny.json, which holds the measured figures and nothing derived: the reductions above are computed from them on every build.

The comparison above is one engine against another on one board. To run the same kind of comparison on your model, see Measure a conversion.