# Calling kernels

After [First kernel](https://ambiqai.github.io/ns-cmsis-nn/getting-started/first-kernel/), most integrations
need more than a tensor pointer: dimensions, operation parameters, workspace,
and sometimes quantization or persistent state. Your application supplies these;
heliaCORE does not load or schedule a model for you.

## Choose an entry point

| Entry point | When to use it | What remains your responsibility |
|---|---|---|
| Wrapper, where provided | Prefer the family wrapper when it supports your operation. For example, `arm_convolve_wrapper_s8` selects among shape-specific convolution implementations. | Match the wrapper's tensor and quantization contracts and use its buffer-size query. |
| Direct kernel | Use when you need a specific documented implementation and can satisfy its constraints. | Check stride, channel, layout, and shape restrictions yourself; use the matching direct-kernel sizing API. |
| Support helper | Use as part of a documented integration, such as preparing convolution weight sums. | Follow its separate layout, lifetime, and return-value contract. |

The compiler selects available scalar, DSP, and MVE code paths. A wrapper can
then choose an implementation from the supplied shape; it is not a runtime CPU
detector. Do not choose a function solely because its name contains `fast`.

## Describe the tensor, not just its size

`cmsis_nn_dims` has `n`, `h`, `w`, and `c` fields, but their meaning is defined
by each function. For `arm_convolve_wrapper_s8`:

| Tensor | Layout | Dimension fields |
|---|---|---|
| Input | NHWC | Batch, height, width, input channels |
| Filter | OHWI | Output channels, filter height, filter width, input channels |
| Output | NHWC | Batch, height, width, output channels |

For a contiguous NHWC tensor, the channel dimension varies fastest. Verify
buffer capacity as well as dimensions, and use the selected operator's layout:
depthwise filters, packed weights, and matrix helpers can use different layouts.
A successful buffer-size query is not full validation of a model's shapes.

## Run a complete pooling example

This example averages a 2×2 input with two interleaved channels into one output
pixel. It uses separate input and output buffers, checks workspace capacity, and
checks both the return status and output values. Build the `pooling` and `nnsupport`
groups with the `s8` type filter, or use a package that includes pooling.

| NHWC input pixels | Per-channel average | Expected output |
|---|---|---|
| `[1, 2]`, `[3, 4]`, `[5, 6]`, `[7, 8]` | `(1+3+5+7)/4`, `(2+4+6+8)/4` | `[4, 5]` |

```c title="pool_example.c"
#include <stdint.h>
#include <string.h>
#include "arm_nnfunctions.h"

int helia_pool_example(void)
{
    const cmsis_nn_dims input_dims = {.n = 1, .h = 2, .w = 2, .c = 2};
    const cmsis_nn_dims filter_dims = {.n = 1, .h = 2, .w = 2, .c = 1};
    const cmsis_nn_dims output_dims = {.n = 1, .h = 1, .w = 1, .c = 2};
    const cmsis_nn_pool_params params = {
        .stride = {.w = 2, .h = 2},
        .padding = {.w = 0, .h = 0},
        .activation = {.min = -128, .max = 127},
    };
    const int8_t input[] = {1, 2, 3, 4, 5, 6, 7, 8};
    int8_t output[2] = {0};
    int32_t scratch[2] = {0};
    const int32_t bytes = arm_avgpool_s8_get_buffer_size(output_dims.w, input_dims.c);
    if (bytes < 0 || (size_t)bytes > sizeof(scratch))
    {
        return 1;
    }
    const cmsis_nn_context ctx = {
        .buf = bytes > 0 ? scratch : NULL,
        .size = bytes,
    };
    const arm_cmsis_nn_status status = arm_avgpool_s8(
        &ctx, &params, &input_dims, input, &filter_dims, &output_dims, output);
    memset(scratch, 0, sizeof(scratch));
    if (status != ARM_CMSIS_NN_SUCCESS)
    {
        return 2;
    }
    return output[0] == 4 && output[1] == 5 ? 0 : 3;
}
```

Calling `helia_pool_example()` returns **0**, with output values **4** and **5**.
Return 1 identifies a workspace sizing failure, 2 a kernel error, and 3 an output
mismatch. These are this example's return codes, not heliaCORE status codes.

The `int32_t` workspace provides the alignment used by this pooling kernel's DSP
accumulator. Its size is specific to these two channels; do not reuse that fixed
size for a different operator or shape. Input and output use the same quantized
scale and zero-point, since this pooling API has no rescaling parameters.

Register the file in your firmware target as shown in
[First kernel](https://ambiqai.github.io/ns-cmsis-nn/getting-started/first-kernel/). Compile it as C; when
calling from C++, declare the function with `extern "C"` in your shared header.

## Prepare convolution state separately

The s8 convolution wrapper accepts both scratch `ctx` and `weight_sum_ctx`.
They are different resources. On MVE, prepare the latter with
`arm_convolve_weight_sum()` using the layer's weights, bias, and input offset.
An allocated but uninitialized weight-sum buffer can produce incorrect output
without an error status.

Pass a non-null weight-sum **context structure** on every build. Non-MVE builds
accept a null buffer inside that context; the weight-sum helper returns
`ARM_CMSIS_NN_NO_IMPL_ERROR` there as its documented no-op case. Do not generalize
that exception to other kernels. See [Memory](https://ambiqai.github.io/ns-cmsis-nn/guide/using-kernels/memory/)
for sizing and reuse, and [Quantization](https://ambiqai.github.io/ns-cmsis-nn/guide/using-kernels/quantization/)
for the parameters.

## Verify before using the output

Check the status returned by the exact API. Not every function returns a status,
and an error code is not a substitute for validating dimensions, capacities, and
parameter ranges before the call. Compare output with a known reference using
representative inputs, including channel counts that are not vector-lane multiples.

For pooling details, use [the pooling APIs](https://ambiqai.github.io/ns-cmsis-nn/reference/api/heliacore/arm_nnfunctions/#arm_avgpool_s8).
For convolution, consult [the kernel index](https://ambiqai.github.io/ns-cmsis-nn/reference/api/heliacore/arm_nnfunctions/#arm_convolve_wrapper_s8)
for the full contract before adapting an upstream example.
