Skip to content
heliaCORE
User guide
HELIA HUB

Acceleration paths

heliaCORE uses Cortex-M compute features available on Ambiq silicon. Your compiler’s target flags determine which implementations can be compiled; each kernel’s data type, shape, and implementation determine which path runs. This applies whether heliaCORE is called by heliaAOT, heliaRT, or your own firmware.

Portable C

The baseline implementation for kernels that provide it. The compiler can still optimize this code for the selected processor.

DSP extensions

Packed integer arithmetic and multiply-accumulate instructions accelerate supported quantized kernels.

Helium / MVE

Vector implementations process multiple tensor elements together, with predication for partial vectors where implemented.

Floating-point support is a separate choice: enable FP16 or FP32 APIs only when your selected kernels, compiler, and target support them. Enabling a data type does not guarantee an MVE implementation for every operation. Check data-type coverage and the kernel reference before choosing an API.

  1. Use the CPU, FPU, and float ABI settings supplied by your board toolchain. See Cortex-M targets.
  2. For source builds, select the operator groups and data types you need in Build configuration. Keep the same target settings on the firmware and kernel sources.
  3. For a prebuilt library, choose an archive with matching settings and enabled features. Changing application flags cannot change the implementation in an already compiled archive. Check its manifest and toolchain compatibility.

The public headers translate compiler feature definitions into kernel guards:

Compiler feature Kernel guard Enables selection of
__ARM_FEATURE_DSP ARM_MATH_DSP DSP implementations where supplied
Integer MVE feature bit ARM_MATH_MVEI Integer vector implementations
Floating-point MVE feature bit ARM_MATH_MVEF, ARM_MATH_MVE_FLOAT16 Floating-point vector implementations, subject to API feature gates

These mappings are defined in arm_nn_math_types.h. ARM_MATH_AUTOVECTORIZE changes explicit vector-path selection in kernels that check it. Use the actual function’s guards when investigating a build rather than assuming one macro describes the whole library.

Inspect the build. Enable verbose output in your build system and find the compile command for the kernel source. Check its CPU/FPU/ABI flags, feature options, and optimization level. With GCC or Clang, reuse that command’s preprocessor settings with -dM -E to inspect the effective macros, or -E to inspect the selected source branch. Remove object-output options when doing so.

Inspect the linked code. Use the matching toolchain’s disassembler on your firmware ELF to inspect the kernel implementation. The link map identifies the object or archive that supplied it. This is especially useful if an SDK already contains a copy of CMSIS-NN or heliaCORE.

Run and measure. First verify output correctness with representative inputs, including shapes that exercise a partial vector. Then measure the call on your Ambiq target. Kernel availability, successful linking, and compiler feature macros do not by themselves establish a speedup.

When performance differs from expectations

Section titled “When performance differs from expectations”
Observation What to check
Expected vector path is absent Effective compiler feature macros, API feature gates, kernel guards, and the linked library
A vector implementation is slower for one shape Tensor sizes, layout, packing/setup costs, memory placement, and the exact function selected
A wrapper selects a different function The wrapper’s shape constraints and buffer requirements in the API reference
Kernel cycles improve but model latency barely changes Time spent in other operators, data movement, and runtime overhead

Use the published kernel measurements as examples for the recorded shapes and configuration. Follow Measurement methodology to compare implementations in your own application.