Calling kernels
After First kernel, most integrations need more than a tensor pointer: dimensions, operation parameters, workspace, and sometimes quantization or persistent state. Your application supplies these; heliaCORE does not load or schedule a model for you.
Choose an entry point
Section titled “Choose an entry point”| Entry point | When to use it | What remains your responsibility |
|---|---|---|
| Wrapper, where provided | Prefer the family wrapper when it supports your operation. For example, arm_convolve_wrapper_s8 selects among shape-specific convolution implementations. |
Match the wrapper’s tensor and quantization contracts and use its buffer-size query. |
| Direct kernel | Use when you need a specific documented implementation and can satisfy its constraints. | Check stride, channel, layout, and shape restrictions yourself; use the matching direct-kernel sizing API. |
| Support helper | Use as part of a documented integration, such as preparing convolution weight sums. | Follow its separate layout, lifetime, and return-value contract. |
The compiler selects available scalar, DSP, and MVE code paths. A wrapper can
then choose an implementation from the supplied shape; it is not a runtime CPU
detector. Do not choose a function solely because its name contains fast.
Describe the tensor, not just its size
Section titled “Describe the tensor, not just its size”cmsis_nn_dims has n, h, w, and c fields, but their meaning is defined
by each function. For arm_convolve_wrapper_s8:
| Tensor | Layout | Dimension fields |
|---|---|---|
| Input | NHWC | Batch, height, width, input channels |
| Filter | OHWI | Output channels, filter height, filter width, input channels |
| Output | NHWC | Batch, height, width, output channels |
For a contiguous NHWC tensor, the channel dimension varies fastest. Verify buffer capacity as well as dimensions, and use the selected operator’s layout: depthwise filters, packed weights, and matrix helpers can use different layouts. A successful buffer-size query is not full validation of a model’s shapes.
Run a complete pooling example
Section titled “Run a complete pooling example”This example averages a 2×2 input with two interleaved channels into one output
pixel. It uses separate input and output buffers, checks workspace capacity, and
checks both the return status and output values. Build the pooling and nnsupport
groups with the s8 type filter, or use a package that includes pooling.
| NHWC input pixels | Per-channel average | Expected output |
|---|---|---|
[1, 2], [3, 4], [5, 6], [7, 8] |
(1+3+5+7)/4, (2+4+6+8)/4 |
[4, 5] |
#include <stdint.h>#include <string.h>#include "arm_nnfunctions.h"
int helia_pool_example(void){ const cmsis_nn_dims input_dims = {.n = 1, .h = 2, .w = 2, .c = 2}; const cmsis_nn_dims filter_dims = {.n = 1, .h = 2, .w = 2, .c = 1}; const cmsis_nn_dims output_dims = {.n = 1, .h = 1, .w = 1, .c = 2}; const cmsis_nn_pool_params params = { .stride = {.w = 2, .h = 2}, .padding = {.w = 0, .h = 0}, .activation = {.min = -128, .max = 127}, }; const int8_t input[] = {1, 2, 3, 4, 5, 6, 7, 8}; int8_t output[2] = {0}; int32_t scratch[2] = {0}; const int32_t bytes = arm_avgpool_s8_get_buffer_size(output_dims.w, input_dims.c); if (bytes < 0 || (size_t)bytes > sizeof(scratch)) { return 1; } const cmsis_nn_context ctx = { .buf = bytes > 0 ? scratch : NULL, .size = bytes, }; const arm_cmsis_nn_status status = arm_avgpool_s8( &ctx, ¶ms, &input_dims, input, &filter_dims, &output_dims, output); memset(scratch, 0, sizeof(scratch)); if (status != ARM_CMSIS_NN_SUCCESS) { return 2; } return output[0] == 4 && output[1] == 5 ? 0 : 3;}The int32_t workspace provides the alignment used by this pooling kernel’s DSP
accumulator. Its size is specific to these two channels; do not reuse that fixed
size for a different operator or shape. Input and output use the same quantized
scale and zero-point, since this pooling API has no rescaling parameters.
Register the file in your firmware target as shown in
First kernel. Compile it as C; when
calling from C++, declare the function with extern "C" in your shared header.
Prepare convolution state separately
Section titled “Prepare convolution state separately”The s8 convolution wrapper accepts both scratch ctx and weight_sum_ctx.
They are different resources. On MVE, prepare the latter with
arm_convolve_weight_sum() using the layer’s weights, bias, and input offset.
An allocated but uninitialized weight-sum buffer can produce incorrect output
without an error status.
Pass a non-null weight-sum context structure on every build. Non-MVE builds
accept a null buffer inside that context; the weight-sum helper returns
ARM_CMSIS_NN_NO_IMPL_ERROR there as its documented no-op case. Do not generalize
that exception to other kernels. See Memory
for sizing and reuse, and Quantization
for the parameters.
Verify before using the output
Section titled “Verify before using the output”Check the status returned by the exact API. Not every function returns a status, and an error code is not a substitute for validating dimensions, capacities, and parameter ranges before the call. Compare output with a known reference using representative inputs, including channel counts that are not vector-lane multiples.
For pooling details, use the pooling APIs. For convolution, consult the kernel index for the full contract before adapting an upstream example.