Kernel benchmarks
heliaCORE puts Arm Helium to work across compute-heavy layers and the smaller operations between them. These measurements compare its MVE, DSP, and portable C (REF) implementations on the same Ambiq Apollo510 EVB.
Faster across the measured workloads
Section titled “Faster across the measured workloads”MVE versus REF
5×
Geometric-mean speedup across the 28 workloads
MVE versus DSP
4×
Geometric-mean speedup across the same workloads
MVE uses fewer cycles
28/28
Than both REF and DSP in this measured set
Equal-weight geometric means across 28 integer workloads. Kernel results, not whole-model speedups. Measurement conditions.
Apollo510 EVB · Cortex-M55 · 96 MHz LP · GCC 14.3.0. Average cycles over 100 calls; REF is compiled portable C and may contain compiler-generated target instructions.
Speedup distribution
All 28 workloads, grouped by MVE speedup versus REF. Bins include their lower bound.
Speedup by workload
Geometric means, 14 workloads per group. Compute includes convolution, matrix math, and pooling; tensor operations include arithmetic and comparisons.
Accelerate more of the inference graph
Section titled “Accelerate more of the inference graph”Two examples from the measured set. Shorter bars mean fewer cycles.
Convolution
90% fewer cycles with MVE than REF
Full s8 convolution, including im2col. The largest speedup in this set.
arm_convolve_s8
32x32x64 k3 oc64 · average cycles per call
Tensor comparisons
70% fewer cycles with MVE than REF
Comparing 4,096 tensor values, beyond the large compute layers.
arm_comparison_s8
n4096 · average cycles per call
Explore operator coverage for indexing, broadcast, reduction, and recurrent APIs beyond this measurement set.
Helium for DSP, ML, and AI workloads
Section titled “Helium for DSP, ML, and AI workloads”MVE brings vector arithmetic to Cortex-M: a single instruction can operate on multiple values, while predication controls which lanes participate. These features serve signal processing and machine learning as well as the neural network kernels measured here.
| Work in an application | Useful MVE capability | Practical benefit |
|---|---|---|
| Filtering, transforms, and feature preparation | Vector arithmetic and multiply-accumulate | Processes groups of samples together in implementations designed for MVE. |
| Quantized and floating-point neural layers | Integer and floating-point vector operations | Accelerates supported arithmetic without changing the model’s chosen numeric format. |
| Tensor operations and irregular lengths | Predicated loads, stores, and arithmetic | Handles partial vectors while keeping useful work in the vector path. |
These are architectural capabilities, not additional heliaCORE benchmark claims. See Arm’s Helium introduction and tail-predication explanation. Use Acceleration to select and verify an implementation. Fewer cycles can free processing time for other work; energy and power benefits require separate measurements.
Correctness before performance
Section titled “Correctness before performance”helia-core-tester supplies generated integer and floating-point test cases, reference-output comparisons, FVP execution, and real-board correctness and performance capture. This lets us examine numeric behavior and execution cost together, rather than treating a successful build as a validated kernel.
The validation guide describes these methods and how to read run evidence. Available test infrastructure does not imply every function, shape, or compiler configuration has passed a hardware run.
See what MVE is doing
Section titled “See what MVE is doing”Timing answers “how long?” PMU counters help explain what happened during the call. helia-core-tester supports selecting named MVE events for hardware capture:
| Question | Counter | What it tells you |
|---|---|---|
| Did the timed code execute vector instructions? | MVE_INST_RETIRED |
Architecturally executed MVE instructions. |
| Is vector work waiting on memory resources? | MVE_STALL_RESOURCE_MEM |
Cycles stalled by MVE memory resource conflicts. |
| Is predication active? | MVE_PRED |
Cycles with one or more predicated beats executing. |
These counters explain a workload; they are not utilization percentages. A high instruction count is not inherently better, and predication events do not report the percentage of active lanes. The cycle dataset above has no accompanying PMU samples. Capture and interpret MVE counters with the tester when investigating your own kernels.
Explore the individual results
Section titled “Explore the individual results”Open a group to inspect and sort the measurements. The same data drives the summary cards, comparisons, and downloadable file.
Download all 28 measurements as CSVConvolution, matrix multiplication, and pooling · 14 workloads
Select a column heading to sort. Ratios above 1 indicate fewer cycles than REF.
arm_convolve_s8contains full im2col | 32x32x64 k3 oc64 | 90,776,688 | 59,320,711 | 7,577,201 | 1.53× | 11.98× |
|---|---|---|---|---|---|---|
arm_convolve_s4contains full im2col | 32x32x64 k3 oc64 | 99,947,570 | 130,619,816 | 18,468,361 | 0.77× | 5.41× |
arm_convolve_s16contains full im2col | 32x32x64 k3 oc64 | 96,570,332 | 82,166,983 | 31,488,224 | 1.18× | 3.07× |
arm_convolve_1x1_s8_fastcontains simplified im2col | 32x32x64 oc64 | 11,885,922 | 8,966,843 | 1,707,351 | 1.33× | 6.96× |
arm_depthwise_conv_s8_optcontains input packing | 32x32x64 k3 | 17,841,554 | 7,509,340 | 1,741,530 | 2.38× | 10.24× |
arm_depthwise_conv_s4_optcontains input packing | 32x32x64 k3 | 7,984,346 | 9,767,607 | 2,188,508 | 0.82× | 3.65× |
arm_depthwise_conv_fast_s16contains input packing | 32x32x64 k3 | 18,069,189 | 7,571,340 | 3,861,216 | 2.39× | 4.68× |
arm_nn_mat_mult_nt_t_s8 | 64x512 × 256x512 | 19,959,308 | 13,725,751 | 2,086,979 | 1.45× | 9.56× |
arm_nn_mat_mult_nt_t_s4 | 64x512 × 256x512 | 26,937,786 | 26,723,602 | 4,197,172 | 1.01× | 6.42× |
arm_nn_vec_mat_mult_t_s8 | 512 × 256 | 491,641 | 331,411 | 69,076 | 1.48× | 7.12× |
arm_nn_vec_mat_mult_t_s16 | 512 × 256 | 562,008 | 421,257 | 99,807 | 1.33× | 5.63× |
arm_nn_vec_mat_mult_t_s4 | 512 × 256 | 504,554 | 588,599 | 88,622 | 0.86× | 5.69× |
arm_avgpool_s8 | 32x32x64 k3 | 13,228,812 | 4,057,857 | 2,135,472 | 3.26× | 6.19× |
arm_avgpool_s16 | 32x32x64 k3 | 6,517,842 | 4,187,804 | 2,408,673 | 1.56× | 2.71× |
Elementwise arithmetic and comparisons · 14 workloads
Select a column heading to sort. Ratios above 1 indicate fewer cycles than REF.
arm_elementwise_add_s8 | n4096 | 250,005 | 285,797 | 61,575 | 0.87× | 4.06× |
|---|---|---|---|---|---|---|
arm_elementwise_add_s16 | n4096 | 252,004 | 252,078 | 53,370 | 1.00× | 4.72× |
arm_elementwise_mul_s8 | n4096 | 98,358 | 105,093 | 22,647 | 0.94× | 4.34× |
arm_elementwise_mul_s16 | n4096 | 77,914 | 79,962 | 28,755 | 0.97× | 2.71× |
arm_elementwise_sub_s8 | n4096 | 254,030 | 285,757 | 61,632 | 0.89× | 4.12× |
arm_elementwise_mul_acc_s16 | n4096 | 86,175 | 90,198 | 31,851 | 0.96× | 2.71× |
arm_elementwise_mul_s16_s8 | n4096 | 86,098 | 84,088 | 28,778 | 1.02× | 2.99× |
arm_elementwise_mul_s16_batch_offset | n4096 | 82,032 | 82,025 | 28,780 | 1.00× | 2.85× |
arm_add_scalar_s8 | n4096 | 131,165 | 179,461 | 42,129 | 0.73× | 3.11× |
arm_sub_scalar_s8 | n4096 | 172,164 | 197,856 | 42,161 | 0.87× | 4.08× |
arm_mul_scalar_s8 | n4096 | 90,169 | 94,308 | 20,087 | 0.96× | 4.49× |
arm_mul_scalar_s16 | n4096 | 77,911 | 77,908 | 27,728 | 1.00× | 2.81× |
arm_comparison_s8 | n4096 | 266,441 | 208,747 | 78,071 | 1.28× | 3.41× |
arm_comparison_s16 | n4096 | 270,764 | 237,720 | 77,986 | 1.14× | 3.47× |
DSP is not faster than REF in every row. Keep the raw cycles, preparation costs, and shape alongside any ratio, especially when choosing a specialized kernel. Follow Measurement to evaluate your own workload. The repository benchmark record is the source of the recorded values; its missing capture details are described in the measurement conditions.