Skip to content
heliaCORE
User guide
HELIA HUB

Calling kernels

After First kernel, most integrations need more than a tensor pointer: dimensions, operation parameters, workspace, and sometimes quantization or persistent state. Your application supplies these; heliaCORE does not load or schedule a model for you.

Entry point When to use it What remains your responsibility
Wrapper, where provided Prefer the family wrapper when it supports your operation. For example, arm_convolve_wrapper_s8 selects among shape-specific convolution implementations. Match the wrapper’s tensor and quantization contracts and use its buffer-size query.
Direct kernel Use when you need a specific documented implementation and can satisfy its constraints. Check stride, channel, layout, and shape restrictions yourself; use the matching direct-kernel sizing API.
Support helper Use as part of a documented integration, such as preparing convolution weight sums. Follow its separate layout, lifetime, and return-value contract.

The compiler selects available scalar, DSP, and MVE code paths. A wrapper can then choose an implementation from the supplied shape; it is not a runtime CPU detector. Do not choose a function solely because its name contains fast.

cmsis_nn_dims has n, h, w, and c fields, but their meaning is defined by each function. For arm_convolve_wrapper_s8:

Tensor Layout Dimension fields
Input NHWC Batch, height, width, input channels
Filter OHWI Output channels, filter height, filter width, input channels
Output NHWC Batch, height, width, output channels

For a contiguous NHWC tensor, the channel dimension varies fastest. Verify buffer capacity as well as dimensions, and use the selected operator’s layout: depthwise filters, packed weights, and matrix helpers can use different layouts. A successful buffer-size query is not full validation of a model’s shapes.

This example averages a 2×2 input with two interleaved channels into one output pixel. It uses separate input and output buffers, checks workspace capacity, and checks both the return status and output values. Build the pooling and nnsupport groups with the s8 type filter, or use a package that includes pooling.

NHWC input pixels Per-channel average Expected output
[1, 2], [3, 4], [5, 6], [7, 8] (1+3+5+7)/4, (2+4+6+8)/4 [4, 5]
pool_example.c
#include <stdint.h>
#include <string.h>
#include "arm_nnfunctions.h"
int helia_pool_example(void)
{
const cmsis_nn_dims input_dims = {.n = 1, .h = 2, .w = 2, .c = 2};
const cmsis_nn_dims filter_dims = {.n = 1, .h = 2, .w = 2, .c = 1};
const cmsis_nn_dims output_dims = {.n = 1, .h = 1, .w = 1, .c = 2};
const cmsis_nn_pool_params params = {
.stride = {.w = 2, .h = 2},
.padding = {.w = 0, .h = 0},
.activation = {.min = -128, .max = 127},
};
const int8_t input[] = {1, 2, 3, 4, 5, 6, 7, 8};
int8_t output[2] = {0};
int32_t scratch[2] = {0};
const int32_t bytes = arm_avgpool_s8_get_buffer_size(output_dims.w, input_dims.c);
if (bytes < 0 || (size_t)bytes > sizeof(scratch))
{
return 1;
}
const cmsis_nn_context ctx = {
.buf = bytes > 0 ? scratch : NULL,
.size = bytes,
};
const arm_cmsis_nn_status status = arm_avgpool_s8(
&ctx, &params, &input_dims, input, &filter_dims, &output_dims, output);
memset(scratch, 0, sizeof(scratch));
if (status != ARM_CMSIS_NN_SUCCESS)
{
return 2;
}
return output[0] == 4 && output[1] == 5 ? 0 : 3;
}

The int32_t workspace provides the alignment used by this pooling kernel’s DSP accumulator. Its size is specific to these two channels; do not reuse that fixed size for a different operator or shape. Input and output use the same quantized scale and zero-point, since this pooling API has no rescaling parameters.

Register the file in your firmware target as shown in First kernel. Compile it as C; when calling from C++, declare the function with extern "C" in your shared header.

The s8 convolution wrapper accepts both scratch ctx and weight_sum_ctx. They are different resources. On MVE, prepare the latter with arm_convolve_weight_sum() using the layer’s weights, bias, and input offset. An allocated but uninitialized weight-sum buffer can produce incorrect output without an error status.

Pass a non-null weight-sum context structure on every build. Non-MVE builds accept a null buffer inside that context; the weight-sum helper returns ARM_CMSIS_NN_NO_IMPL_ERROR there as its documented no-op case. Do not generalize that exception to other kernels. See Memory for sizing and reuse, and Quantization for the parameters.

Check the status returned by the exact API. Not every function returns a status, and an error code is not a substitute for validating dimensions, capacities, and parameter ranges before the call. Compare output with a known reference using representative inputs, including channel counts that are not vector-lane multiples.

For pooling details, use the pooling APIs. For convolution, consult the kernel index for the full contract before adapting an upstream example.