Skip to content
heliaCORE
User guide
HELIA HUB

Operator coverage

heliaCORE builds on Arm CMSIS-NN with additional operators and implementations for production models on Ambiq silicon. Its coverage includes convolution and dense layers, tensor indexing and movement, reductions, comparisons, activations, and recurrent operations, with both quantized and floating-point APIs.

Start with the operation your model needs, then check the exact data type and shape in the kernel index. Search by name, filter by family or data type, and open the function for its contract.

The links below open the corresponding API family. Each family contains multiple operations and variants; a family containing FP16 functions does not imply FP16 support for every operation in that family.

FamilyOperations and variantsNumeric tags present
ConvolutionConv2D, depthwise, transpose convolution, wrappers, and buffer helpers.s4, s8, s16, f16, f32
Fully connectedDense layers, batch matmul paths, and scratch sizing helpers.s4, s8, s16, f16, f32
ElementwiseAdd, sub, mul, square difference, min/max, batch norm, select, and arithmetic glue.s8, s16, f16, f32
Reduction and comparisonArgmin/argmax, min/max reductions, comparisons, means, vector sums, and where.s8, s16, s64, f16, f32
ActivationReLU, LeakyReLU, PReLU, Hard-Swish, Logistic, Tanh, and clamp.s8, s16, q7, q15, f16, f32
Data movementPad, reshape, transpose, concatenate, gather, resize, strided slice, tile, broadcast, scatter, mirror pad, and sequence/slice-update utilities.s8, s16, s32, f16, f32
Pooling, softmax, quantizationClassifier tail APIs plus dtype conversion and requantization utilities.s8, s16, u8, f16, f32
SequenceLSTM, GRU, and SVDF functions for temporal workloads.s8, s16, f16, f32

The table is generated from declarations in the public kernel headers using the same extracted API model as the reference. It includes wrapper and buffer-sizing functions. It is an API inventory, not a count of distinct model operators or accelerated implementations. Detailed function counts explain the counting rules.

Format Typical API tag Check before using
A8W8 s8 Activation and weight quantization, offsets, bias type, and per-channel parameters
A16W8 s16 on the relevant compute APIs Weight and accumulator types in the signature; s16 alone does not describe all arguments
4-bit weights s4 Packed weight layout and supported shapes; this is not a general 4-bit activation format
FP16 f16 Opt-in float API and build support, target features, and function-specific constraints
FP32 f32 Opt-in float API and the selected target implementation

The inherited q7 and q15 names use separate fixed-point conventions. Do not choose them solely because their storage width matches a quantized tensor. See Data types and Quantization before preparing inputs.

Before treating an operator as supported in your application, confirm all four:

  1. Contract: the function accepts the tensor layout, shape, strides, padding, and numeric format required by your model.
  2. Build: its source group, support functions, and data type are included in your library, and public headers have matching feature definitions.
  3. Memory: output, state, and scratch buffers satisfy the function’s requirements.
  4. Execution: the selected path produces the expected result for representative inputs on your target. DSP or Helium coverage can depend on shape and format.

Use Calling kernels for the call sequence and Acceleration for implementation selection.

Arm CMSIS-NN supplies the foundation. heliaCORE extends it with operations such as gather, scatter, graph updates, reductions, and floating-point arithmetic for Ambiq applications. The CMSIS-NN relationship gives concrete additions and the source revisions used for comparison.