# heliaCORE

Neural network kernels for Ambiq silicon.

heliaCORE is a neural network kernel library optimized for Ambiq silicon.
It extends Arm CMSIS-NN with broader operator coverage, quantized and floating-point
kernels, and Cortex-M DSP and Helium implementations. Use it through heliaAOT or
heliaRT, or integrate kernels directly into your firmware.

[Choose your inference path](https://ambiqai.github.io/ns-cmsis-nn/#inference)
[Explore the kernels](https://ambiqai.github.io/ns-cmsis-nn/reference/kernel-index/)
v

01
Beyond convolutionTensor indexing, reductions, broadcast math and recurrent layers.

02
DSP + HeliumOptimized Cortex-M compute paths.A8W8A16W8FP16 / FP32

03
One build foundationCMake · CMSIS-Pack · Zephyr · neuralSPOT-X

Check coverage by operator and target ↗

Kernel catalog ↗
DSP + Helium ↗
Data-type coverage ↗
Build integration ↗

Optimized computation throughout your model
Production models do more than convolution and matrix multiplication. They
also move tensors, combine results, and maintain state. heliaCORE gives model
integrations optimized library kernels for more of these operations, including
the work between convolution and dense layers.
For most applications, start with heliaAOT or heliaRT,
which use heliaCORE to execute supported model operations on Ambiq devices.
heliaAOT is the recommended starting point for latency, power, and memory
efficiency. Direct kernel integration is available for custom runtimes and
firmware that need control over individual operations.

Operator coverage
Coverage across production model graphs

Alongside the convolution, dense, and other kernels inherited from Arm CMSIS-NN,
heliaCORE adds extensive support for tensor indexing, reductions, comparisons,
broadcast math, and recurrent networks. This extends kernel coverage to more of
the operations that make up a complete model. Your inference runtime or compiler
also determines which model operations it can map to these kernels.

Tensor indexing and updates
Gather, GatherND, ScatterND, Tile, Select, ReverseSequence, and DynamicUpdateSlice extend the graph operations available to your model.

Reductions and comparisons
Reduce sum, minimum and maximum, ArgMin, ArgMax, and comparison kernels cover decisions and aggregation within the graph.

More floating-point operations
FP16 and FP32 additions include broadcast add, subtract and multiply, gather, and reductions, alongside convolution and dense layers.

Extended recurrent support
FP16 and FP32 GRU kernels expand recurrent model support beyond the inherited LSTM and SVDF implementations.

[Explore operator coverage](https://ambiqai.github.io/ns-cmsis-nn/guide/coverage/operator-coverage/)
[Explore the CMSIS-NN foundation and extensions](https://ambiqai.github.io/ns-cmsis-nn/guide/coverage/compared-with-cmsis-nn/)

Cortex-M acceleration
Put DSP and Helium to work

Use the compute features already in the processor. heliaCORE selects implementations
from the target's compiler flags, with specialized paths for supported kernels
and portable C implementations where the selected API provides one.

DSP extensions
Packed integer computation
DSP instructions process packed integer values and accelerate multiply-accumulate
operations used by quantized kernels.

Helium / MVE
Vectorized neural network math
Hand-tuned vector implementations accelerate integer and floating-point operations,
with predicated processing for the ends of tensors.

Memory control
Buffers owned by your firmware
No dynamic allocation inside the library. Query scratch-buffer requirements
and supply the memory from your application's allocation strategy.

[Explore acceleration paths](https://ambiqai.github.io/ns-cmsis-nn/guide/architecture/acceleration-paths/)
[View kernel benchmarks](https://ambiqai.github.io/ns-cmsis-nn/guide/performance/kernel-benchmarks/)
[Supported targets](https://ambiqai.github.io/ns-cmsis-nn/guide/architecture/cortex-m-targets/)
[Toolchain requirements](https://ambiqai.github.io/ns-cmsis-nn/guide/architecture/toolchains/)

Build integration
One kernel library. Your build system.

Add heliaCORE to an existing firmware project or use it through the HELIA stack.
Choose the integration that matches your tools, then configure the operator groups
and data types your application needs.

Build only what your application needs

One CMake manifest defines the kernel sources for standalone, Zephyr, and
neuralSPOT-X builds. Conditional source selection lets you leave unused operator
groups and optional floating-point kernels out of the build.

One shared source manifest
cmake/ns_cmsis_nn.cmake

Choose operator groups
Example configuration

ConvolutionON
Fully connectedON
LSTMOFF
SVDFOFF

Enable the data types you need
Filter sources by data type. FP16 and FP32 are separate, opt-in build options.

FP16 kernelsOFF
FP32 kernelsOFF

Floating-point defaults shown

Selected sources compile into your library.
Disabled groups and floating-point variants stay out of the source build.
Selection happens at build time, with no runtime switch.

CMSIS-Pack source lists are checked against the repository to keep packaged
kernels aligned with source builds.

- [CMake](https://ambiqai.github.io/ns-cmsis-nn/getting-started/cmake/): For an existing CMake application. Install a prebuilt library package and link the exported target, or build from source with the groups and data types you select.
- [CMSIS-Pack](https://ambiqai.github.io/ns-cmsis-nn/getting-started/cmsis-pack/): For projects using CMSIS tooling. Add the Ambiq pack and choose source components or a prebuilt library for your target.
- [Zephyr](https://ambiqai.github.io/ns-cmsis-nn/getting-started/zephyr/): For Zephyr applications. Add the module to your workspace and configure kernel support through Kconfig, using sources or a prebuilt archive.
- [neuralSPOT-X](https://ambiqai.github.io/ns-cmsis-nn/getting-started/neuralspot-x/): For applications built with Ambiq’s development SDK. Bring heliaCORE into the SDK’s CMake build, including applications that use heliaRT.
[Choose your integration](https://ambiqai.github.io/ns-cmsis-nn/getting-started/)
[Configure kernels and data types](https://ambiqai.github.io/ns-cmsis-nn/guide/architecture/build-path-selection/)

Powered by heliaCORE
Choose your inference integration

Choose heliaAOT to compile your model ahead of time, or heliaRT for a LiteRT for
Microcontrollers runtime integration. Both use heliaCORE kernels for supported
operations. Their guides cover model compatibility and deployment.

- [heliaAOT](https://ambiqai.github.io/helia-aot/): Our ahead-of-time inference path, recommended for latency, power, and memory efficiency. Compile your model into standalone C inference code that uses heliaCORE kernels, with memory planned before deployment.
- [heliaRT](https://ambiqai.github.io/helia-rt/): Our optimized fork of LiteRT for Microcontrollers. Run models through a familiar runtime integration, powered by heliaCORE’s expanded kernel coverage and Ambiq-targeted optimizations.
Building a custom runtime or integrating kernels directly?
[Get started with direct integration](https://ambiqai.github.io/ns-cmsis-nn/getting-started/)
