Skip to content
heliaCORE
User guide
HELIA HUB

Measurement methodology

Measure the kernel in the configuration your application uses. Compare output correctness first, then execution cost. A faster kernel is useful only if it preserves the required behavior for your tensors.

The kernel benchmark tables retain the following conditions from the repository measurement record.

Parameter Recorded value
Board Apollo510 EVB (Cortex-M55)
Clock 96 MHz (LP mode)
Counter DWT CYCCNT
Implementations Portable C (REF), DSP, MVE / Helium
Compiler arm-none-eabi-gcc 14.3.0
Optimization -O3, CMSIS_NN_USE_REQUANTIZE_INLINE_ASSEMBLY
Iterations 100 per kernel
Metric Average CPU cycles per call

All three implementations were measured on the same SoC at the same clock. REF is compiled portable C, which the compiler may optimize with target-specific instructions. Ratios divide one implementation’s average cycles by another’s; the headline geometric means give each listed workload equal weight. The detailed tables retain individual ratios. For ratios r₁ through rₙ, the geometric mean is exp(mean(log(r))). It does not sum timings from unrelated shapes or represent the latency of a complete model.

Record the kernel symbol, tensor dimensions and layout, stride/padding where applicable, data types, quantization parameters, and scratch-buffer size. Preserve the inputs and expected outputs so the comparison can be repeated. Include shapes representative of your model, rather than choosing only a favorable vector-aligned size.

2. Record the build and board configuration

Section titled “2. Record the build and board configuration”

Keep the heliaCORE commit or release, compiler version, full compile/link flags, feature definitions, firmware ELF, and link map with the results. Record the board and clock configuration, code/data/scratch placement, and relevant cache settings. When comparing implementations, hold these conditions fixed and record the source or build changes used to select each path.

See Acceleration paths for ways to verify the selected implementation. A REF label alone does not prove which instructions the compiler emitted.

Run the kernel and compare its result with expected output under the API’s numeric contract. Check return status where provided. Reinitialize any inputs that the kernel modifies before each iteration, so repeated calls perform the same work. Keep validation and logging outside the timed interval.

Use the board SDK’s supported timing facilities. If you use DWT CYCCNT, verify that the counter is enabled and advancing on your target, account for counter wraparound, and measure the overhead of the timing reads.

State whether allocation, input packing, weight preparation, and cache warmup are inside or outside the interval. Measure the full operation your application pays for; if preparation is reused across calls, report it separately. Choose and record an interrupt policy that represents your comparison, rather than silently mixing interrupted and uninterrupted runs.

Record the iteration count, individual samples where practical, and summary statistics such as minimum, median, and mean. Explain warmup and any excluded samples. Report the actual cycle counts alongside the ratio:

speedup = baseline average cycles / candidate average cycles

A value greater than 1 means fewer cycles for the candidate; below 1 means more. Convert cycles to elapsed time only with the measured clock configuration. Cycle counts alone do not measure energy or establish battery-life improvements.

Use helia-core-tester’s hardware runner to combine correctness checks with timing and named PMU events. Complete that repository’s hardware/toolchain setup first. From its checkout, this example builds, flashes when required, and runs convolution cases on a connected board:

Terminal window
uv run helia_core_tester hardware run \
--board apollo510_evb \
--suite int \
--family ConvolutionFunctions \
--pmu-counters mve:ARM_PMU_MVE_INST_RETIRED,ARM_PMU_MVE_STALL_RESOURCE_MEM,ARM_PMU_MVE_PRED

Select the board and probe explicitly when more than one is connected. The tester README documents device selection, firmware setup, counter support and limits. The example is a capture workflow, not a command to reproduce the older cycle tables: those tables do not record a matching source revision or capture plan.

Event Interpretation
ARM_PMU_MVE_INST_RETIRED MVE instructions architecturally executed. Check that vector code runs in the measured window.
ARM_PMU_MVE_FP_MAC_RETIRED Floating-point MVE multiply or multiply-accumulate instructions executed. Useful when examining a supported float compute kernel.
ARM_PMU_MVE_STALL_RESOURCE_MEM Stalls attributed to MVE memory-resource conflicts. Investigate layout, access patterns, and placement alongside cycle counts.
ARM_PMU_MVE_PRED Cycles with one or more predicated beats executing. This does not measure lane utilization.

Definitions come from CMSIS PMU events. Availability depends on the target and capture configuration; an unsupported counter is not a zero measurement. The tester’s event catalog lists selectable names.

  • Check output correctness first. Then inspect overflow_detected, valid_for_regression, and the individual counter support flags. The validity field does not replace a correctness result.
  • case_summary.csv reports counter medians per invocation. Raw samples retain the capture values and iteration counts; apply the documented normalization before comparing them with the summary.
  • Selecting more counters can require multiple passes. Each pass replays the workload, so keep input and state preparation consistent. Do not add replayed cycle counts into a single-call latency or assume all events were simultaneous.
  • PMU cycles and DWT cycles can use different capture boundaries. Compare each against its own window and overhead rather than forcing their values to match.
  • Retired instruction counts and stall events explain activity. They are not direct measurements of throughput, vector utilization, power, or energy.

Keep the raw samples, case summary, session manifest, correctness results and firmware artifacts together. See Validation for the evidence needed to support a result.

Kernel timings isolate one operation and shape. Model latency also includes other operators, tensor movement, preparation, and runtime overhead. After improving a kernel, measure end-to-end inference in the intended heliaAOT, heliaRT, or custom-runtime application to establish the application benefit.