Skip to content
heliaCORE
Home
HELIA HUB

heliaCORE

Neural network kernels for Ambiq silicon.

heliaCORE is a neural network kernel library optimized for Ambiq silicon. It extends Arm CMSIS-NN with broader operator coverage, quantized and floating-point kernels, and Cortex-M DSP and Helium implementations. Use it through heliaAOT or heliaRT, or integrate kernels directly into your firmware.

Beyond convolution

Tensor indexing, reductions, broadcast math and recurrent layers.

DSP + Helium

Optimized Cortex-M compute paths.

A8W8A16W8FP16 / FP32
One build foundation

CMake · CMSIS-Pack · Zephyr · neuralSPOT-X

Check coverage by operator and target

Optimized computation throughout your model

Production models do more than convolution and matrix multiplication. They also move tensors, combine results, and maintain state. heliaCORE gives model integrations optimized library kernels for more of these operations, including the work between convolution and dense layers.

For most applications, start with heliaAOT or heliaRT, which use heliaCORE to execute supported model operations on Ambiq devices. heliaAOT is the recommended starting point for latency, power, and memory efficiency. Direct kernel integration is available for custom runtimes and firmware that need control over individual operations.

Operator coverage

Coverage across production model graphs

Alongside the convolution, dense, and other kernels inherited from Arm CMSIS-NN, heliaCORE adds extensive support for tensor indexing, reductions, comparisons, broadcast math, and recurrent networks. This extends kernel coverage to more of the operations that make up a complete model. Your inference runtime or compiler also determines which model operations it can map to these kernels.

Tensor indexing and updates

Gather, GatherND, ScatterND, Tile, Select, ReverseSequence, and DynamicUpdateSlice extend the graph operations available to your model.

Reductions and comparisons

Reduce sum, minimum and maximum, ArgMin, ArgMax, and comparison kernels cover decisions and aggregation within the graph.

More floating-point operations

FP16 and FP32 additions include broadcast add, subtract and multiply, gather, and reductions, alongside convolution and dense layers.

Extended recurrent support

FP16 and FP32 GRU kernels expand recurrent model support beyond the inherited LSTM and SVDF implementations.

Cortex-M acceleration

Put DSP and Helium to work

Use the compute features already in the processor. heliaCORE selects implementations from the target’s compiler flags, with specialized paths for supported kernels and portable C implementations where the selected API provides one.

DSP extensions

Packed integer computation

DSP instructions process packed integer values and accelerate multiply-accumulate operations used by quantized kernels.

Helium / MVE

Vectorized neural network math

Hand-tuned vector implementations accelerate integer and floating-point operations, with predicated processing for the ends of tensors.

Memory control

Buffers owned by your firmware

No dynamic allocation inside the library. Query scratch-buffer requirements and supply the memory from your application’s allocation strategy.

Build integration

One kernel library. Your build system.

Add heliaCORE to an existing firmware project or use it through the HELIA stack. Choose the integration that matches your tools, then configure the operator groups and data types your application needs.

CMake single source of truth

Build only what your application needs

One CMake manifest defines the kernel sources for standalone, Zephyr, and neuralSPOT-X builds. Conditional source selection lets you leave unused operator groups and optional floating-point kernels out of the build.

One shared source manifestcmake/ns_cmsis_nn.cmake

Choose operator groups

Example configuration

Convolution
ON
Fully connected
ON
LSTM
OFF
SVDF
OFF

Enable the data types you need

Filter sources by data type. FP16 and FP32 are separate, opt-in build options.

FP16 kernels
OFF
FP32 kernels
OFF

Floating-point defaults shown

Selected sources compile into your library.

Disabled groups and floating-point variants stay out of the source build. Selection happens at build time, with no runtime switch.

CMSIS-Pack source lists are checked against the repository to keep packaged kernels aligned with source builds.

Powered by heliaCORE

Choose your inference integration

Choose heliaAOT to compile your model ahead of time, or heliaRT for a LiteRT for Microcontrollers runtime integration. Both use heliaCORE kernels for supported operations. Their guides cover model compatibility and deployment.