Skip to content
heliaCORE
User guide
HELIA HUB

Kernel benchmarks

heliaCORE puts Arm Helium to work across compute-heavy layers and the smaller operations between them. These measurements compare its MVE, DSP, and portable C (REF) implementations on the same Ambiq Apollo510 EVB.

MVE versus REF

5×

Geometric-mean speedup across the 28 workloads

MVE versus DSP

4×

Geometric-mean speedup across the same workloads

MVE uses fewer cycles

28/28

Than both REF and DSP in this measured set

Equal-weight geometric means across 28 integer workloads. Kernel results, not whole-model speedups. Measurement conditions.

Apollo510 EVB · Cortex-M55 · 96 MHz LP · GCC 14.3.0. Average cycles over 100 calls; REF is compiled portable C and may contain compiler-generated target instructions.

Speedup distribution

Number of workloads<2×: 0 workloads0<2×2–4×: 11 workloads112–4×4–6×: 10 workloads104–6×6–8×: 4 workloads46–8×≥8×: 3 workloads3≥8×

All 28 workloads, grouped by MVE speedup versus REF. Bins include their lower bound.

Speedup by workload

Compute layersCompute layers: 5.86× MVE speedup versus REFvs REF6×Compute layers: 4.20× MVE speedup versus DSPvs DSP4×Tensor operationsTensor operations: 3.50× MVE speedup versus REFvs REF3×Tensor operations: 3.62× MVE speedup versus DSPvs DSP4×0×2×4×6×

Geometric means, 14 workloads per group. Compute includes convolution, matrix math, and pooling; tensor operations include arithmetic and comparisons.

Two examples from the measured set. Shorter bars mean fewer cycles.

Convolution

90% fewer cycles with MVE than REF

REF90,776,688
DSP59,320,711
MVE7,577,201

Full s8 convolution, including im2col. The largest speedup in this set.

arm_convolve_s8
32x32x64 k3 oc64 · average cycles per call

Tensor comparisons

70% fewer cycles with MVE than REF

REF266,441
DSP208,747
MVE78,071

Comparing 4,096 tensor values, beyond the large compute layers.

arm_comparison_s8
n4096 · average cycles per call

Explore operator coverage for indexing, broadcast, reduction, and recurrent APIs beyond this measurement set.

MVE brings vector arithmetic to Cortex-M: a single instruction can operate on multiple values, while predication controls which lanes participate. These features serve signal processing and machine learning as well as the neural network kernels measured here.

Work in an application Useful MVE capability Practical benefit
Filtering, transforms, and feature preparation Vector arithmetic and multiply-accumulate Processes groups of samples together in implementations designed for MVE.
Quantized and floating-point neural layers Integer and floating-point vector operations Accelerates supported arithmetic without changing the model’s chosen numeric format.
Tensor operations and irregular lengths Predicated loads, stores, and arithmetic Handles partial vectors while keeping useful work in the vector path.

These are architectural capabilities, not additional heliaCORE benchmark claims. See Arm’s Helium introduction and tail-predication explanation. Use Acceleration to select and verify an implementation. Fewer cycles can free processing time for other work; energy and power benefits require separate measurements.

helia-core-tester supplies generated integer and floating-point test cases, reference-output comparisons, FVP execution, and real-board correctness and performance capture. This lets us examine numeric behavior and execution cost together, rather than treating a successful build as a validated kernel.

The validation guide describes these methods and how to read run evidence. Available test infrastructure does not imply every function, shape, or compiler configuration has passed a hardware run.

Timing answers “how long?” PMU counters help explain what happened during the call. helia-core-tester supports selecting named MVE events for hardware capture:

Question Counter What it tells you
Did the timed code execute vector instructions? MVE_INST_RETIRED Architecturally executed MVE instructions.
Is vector work waiting on memory resources? MVE_STALL_RESOURCE_MEM Cycles stalled by MVE memory resource conflicts.
Is predication active? MVE_PRED Cycles with one or more predicated beats executing.

These counters explain a workload; they are not utilization percentages. A high instruction count is not inherently better, and predication events do not report the percentage of active lanes. The cycle dataset above has no accompanying PMU samples. Capture and interpret MVE counters with the tester when investigating your own kernels.

Open a group to inspect and sort the measurements. The same data drives the summary cards, comparisons, and downloadable file.

Download all 28 measurements as CSV
Convolution, matrix multiplication, and pooling · 14 workloads

Select a column heading to sort. Ratios above 1 indicate fewer cycles than REF.

Convolution, matrix multiplication, and pooling results
arm_convolve_s8contains full im2col32x32x64 k3 oc6490,776,68859,320,7117,577,2011.53×11.98×
arm_convolve_s4contains full im2col32x32x64 k3 oc6499,947,570130,619,81618,468,3610.77×5.41×
arm_convolve_s16contains full im2col32x32x64 k3 oc6496,570,33282,166,98331,488,2241.18×3.07×
arm_convolve_1x1_s8_fastcontains simplified im2col32x32x64 oc6411,885,9228,966,8431,707,3511.33×6.96×
arm_depthwise_conv_s8_optcontains input packing32x32x64 k317,841,5547,509,3401,741,5302.38×10.24×
arm_depthwise_conv_s4_optcontains input packing32x32x64 k37,984,3469,767,6072,188,5080.82×3.65×
arm_depthwise_conv_fast_s16contains input packing32x32x64 k318,069,1897,571,3403,861,2162.39×4.68×
arm_nn_mat_mult_nt_t_s864x512 × 256x51219,959,30813,725,7512,086,9791.45×9.56×
arm_nn_mat_mult_nt_t_s464x512 × 256x51226,937,78626,723,6024,197,1721.01×6.42×
arm_nn_vec_mat_mult_t_s8512 × 256491,641331,41169,0761.48×7.12×
arm_nn_vec_mat_mult_t_s16512 × 256562,008421,25799,8071.33×5.63×
arm_nn_vec_mat_mult_t_s4512 × 256504,554588,59988,6220.86×5.69×
arm_avgpool_s832x32x64 k313,228,8124,057,8572,135,4723.26×6.19×
arm_avgpool_s1632x32x64 k36,517,8424,187,8042,408,6731.56×2.71×

Elementwise arithmetic and comparisons · 14 workloads

Select a column heading to sort. Ratios above 1 indicate fewer cycles than REF.

Elementwise and comparison results
arm_elementwise_add_s8n4096250,005285,79761,5750.87×4.06×
arm_elementwise_add_s16n4096252,004252,07853,3701.00×4.72×
arm_elementwise_mul_s8n409698,358105,09322,6470.94×4.34×
arm_elementwise_mul_s16n409677,91479,96228,7550.97×2.71×
arm_elementwise_sub_s8n4096254,030285,75761,6320.89×4.12×
arm_elementwise_mul_acc_s16n409686,17590,19831,8510.96×2.71×
arm_elementwise_mul_s16_s8n409686,09884,08828,7781.02×2.99×
arm_elementwise_mul_s16_batch_offsetn409682,03282,02528,7801.00×2.85×
arm_add_scalar_s8n4096131,165179,46142,1290.73×3.11×
arm_sub_scalar_s8n4096172,164197,85642,1610.87×4.08×
arm_mul_scalar_s8n409690,16994,30820,0870.96×4.49×
arm_mul_scalar_s16n409677,91177,90827,7281.00×2.81×
arm_comparison_s8n4096266,441208,74778,0711.28×3.41×
arm_comparison_s16n4096270,764237,72077,9861.14×3.47×

DSP is not faster than REF in every row. Keep the raw cycles, preparation costs, and shape alongside any ratio, especially when choosing a specialized kernel. Follow Measurement to evaluate your own workload. The repository benchmark record is the source of the recorded values; its missing capture details are described in the measurement conditions.