Acceleration paths
heliaCORE uses Cortex-M compute features available on Ambiq silicon. Your compiler’s target flags determine which implementations can be compiled; each kernel’s data type, shape, and implementation determine which path runs. This applies whether heliaCORE is called by heliaAOT, heliaRT, or your own firmware.
Available implementations
Section titled “Available implementations”Portable C
DSP extensions
Helium / MVE
Floating-point support is a separate choice: enable FP16 or FP32 APIs only when your selected kernels, compiler, and target support them. Enabling a data type does not guarantee an MVE implementation for every operation. Check data-type coverage and the kernel reference before choosing an API.
Configure your build
Section titled “Configure your build”- Use the CPU, FPU, and float ABI settings supplied by your board toolchain. See Cortex-M targets.
- For source builds, select the operator groups and data types you need in Build configuration. Keep the same target settings on the firmware and kernel sources.
- For a prebuilt library, choose an archive with matching settings and enabled features. Changing application flags cannot change the implementation in an already compiled archive. Check its manifest and toolchain compatibility.
The public headers translate compiler feature definitions into kernel guards:
| Compiler feature | Kernel guard | Enables selection of |
|---|---|---|
__ARM_FEATURE_DSP |
ARM_MATH_DSP |
DSP implementations where supplied |
| Integer MVE feature bit | ARM_MATH_MVEI |
Integer vector implementations |
| Floating-point MVE feature bit | ARM_MATH_MVEF, ARM_MATH_MVE_FLOAT16 |
Floating-point vector implementations, subject to API feature gates |
These mappings are defined in
arm_nn_math_types.h.
ARM_MATH_AUTOVECTORIZE changes explicit vector-path selection in kernels that
check it. Use the actual function’s guards when investigating a build rather
than assuming one macro describes the whole library.
Verify the implementation
Section titled “Verify the implementation”Inspect the build. Enable verbose output in your build system and find the
compile command for the kernel source. Check its CPU/FPU/ABI flags, feature
options, and optimization level. With GCC or Clang, reuse that command’s
preprocessor settings with -dM -E to inspect the effective macros, or -E to
inspect the selected source branch. Remove object-output options when doing so.
Inspect the linked code. Use the matching toolchain’s disassembler on your firmware ELF to inspect the kernel implementation. The link map identifies the object or archive that supplied it. This is especially useful if an SDK already contains a copy of CMSIS-NN or heliaCORE.
Run and measure. First verify output correctness with representative inputs, including shapes that exercise a partial vector. Then measure the call on your Ambiq target. Kernel availability, successful linking, and compiler feature macros do not by themselves establish a speedup.
When performance differs from expectations
Section titled “When performance differs from expectations”| Observation | What to check |
|---|---|
| Expected vector path is absent | Effective compiler feature macros, API feature gates, kernel guards, and the linked library |
| A vector implementation is slower for one shape | Tensor sizes, layout, packing/setup costs, memory placement, and the exact function selected |
| A wrapper selects a different function | The wrapper’s shape constraints and buffer requirements in the API reference |
| Kernel cycles improve but model latency barely changes | Time spent in other operators, data movement, and runtime overhead |
Use the published kernel measurements as examples for the recorded shapes and configuration. Follow Measurement methodology to compare implementations in your own application.