# heliaCORE
---
# heliaCORE.Acti
Perform activation layers, including ReLU (Rectified Linear Unit), sigmoid and tanh
## arm_nn_activation_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_activation_f32(
const float32_t *input,
float32_t *output,
int32_t size,
arm_nn_activation_type_flt type,
float32_t act_param
)
```
Elementwise activation.
:::note
The RELU, RELU6 and LEAKY_RELU legs classify NaN on the integer bit pattern (#380 / #382), so a NaN input comes back as NaN at every optimization level on the gated toolchains, including the shipped -Ofast. This holds on both the scalar and the MVE (cortex-m55) build paths; the MVE RELU/RELU6 legs restore the NaN lanes that vmaxnmq/vminnmq suppress. The MVE TANH leg returns a NaN input unchanged and keeps the sign of zero, also decided on the bit pattern (#635). The scalar TANH leg returns NaN where there is no hardware floating point (__ARM_FP undefined, e.g. Cortex-M0), where it too classifies NaN on the bit pattern (quieting a signalling NaN), and elsewhere only in builds without -ffinite-math-only. SIGMOID and HARDSWISH are outside this contract; see the per-helper notes in `Include/Internal/arm_nn_activation_flt.h`.
:::
:::note
The HARDSWISH leg's scalar helper (arm_nn_hardswish_scalar_f32, serving every build that does not take the MVE float path no MVE float support, or MVE present but not used, e.g. under ARM_MATH_AUTOVECTORIZE) keeps the legacy separately rounded multiply-and-add gate and can differ by an ulp in the curved region from the standalone `arm_hard_swish_f32`, whose gate is a correctly rounded fma; the mux's MVE helper (arm_nn_vhardswish_mve_f32) uses vfmaq and agrees with that kernel. Callers that need bit-exact, leg-agreeing hard swish or the documented NaN/Inf contract should call `arm_hard_swish_f32` directly.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input samples. |
| output | float32_t * | out | Pointer to the output samples. |
| size | int32_t | in | Number of elements to process. |
| type | arm_nn_activation_type_flt | in | Activation selector. |
| act_param | float32_t | in | Extra activation parameter. Used for parameterized activations such as leaky ReLU. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:609`
## arm_prelu_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_prelu_f32(
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *alpha_dims,
const float32_t *alpha,
const cmsis_nn_dims *output_dims,
float32_t *output
)
```
Parametric ReLU for float32 data.
Computes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Must equal output_dims. |
| input | const float32_t * | in | Pointer to the input tensor. |
| alpha_dims | const cmsis_nn_dims * | in | Alpha tensor dimensions. |
| alpha | const float32_t * | in | Pointer to the alpha (slope) tensor. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:631`
## arm_hard_swish_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_hard_swish_f32(const float32_t *input, float32_t *output, int32_t size)
```
Hard swish activation for float32 data.
Computes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise, evaluated as x * clamp(fma(x, 1/6, 0.5), 0, 1) so the saturated regions are exact: x >= 3 returns x bit-exactly and x <= -3 returns zero exactly (a negative zero, as IEEE negative * +0.0). In the curved region -3 < x < 3 the gate is a correctly rounded fused multiply-add on both build paths, so the scalar and MVE (cortex-m55) legs agree bit-exactly on every numeric normal input. Two carve-outs, both rooted in Armv8.1-M MVE floating-point arithmetic using the architecture's Standard FPSCR value DN=1 and FZ=1 hard-wired, FZ16 passed through (Arm v8-M ARM, DDI 0553B.l, StandardFPSCRValue(), selected by the MVE FP pseudocode's fpscr_controlled=FALSE): NaN lanes agree in NaN-ness but not necessarily in payload (forced DN makes the MVE leg canonicalize payloads the scalar leg preserves), and the MVE leg flushes f32 subnormal operands and results to a signed zero regardless of FPSCR.FZ, where the scalar leg with FZ clear keeps them. Both reference models (FVP Corstone-300 and QEMU mps3-an547) exhibit the flush identically; it has not been executed on silicon, where the same architectural behavior is required. Near the lower knot the absolute contract is the meaningful one: for x just above -3 the output error is dominated by the gate constant's representation error, bounded by |x^2 * (1/6f - 1/6)| ~ 4.5e-08 near x = -3 (e.g. nextafterf(-3, 0) returns -7.45e-08 against a float64 -1.19e-07 millions of ulps of the tiny result, well inside the 1e-6 absolute contract), and where the gate underflows to exactly zero the kernel returns -0.0 with unbounded relative error. In-place operation (output == input) is supported on both legs; each element is read before it is written.
:::note
NaN propagates (TensorFlow Lite semantics): a NaN input element yields NaN at that output element at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast: propagation rides the final multiply x * gate a NaN x makes the product NaN whatever the gate resolved to rather than a compare-and-select that -ffinite-math-only could fold. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. +Inf returns +Inf (the gate is 1). -Inf returns NaN, not the mathematical limit 0: the gate is 0 there and (-Inf) * 0 is NaN by IEEE 754, the same result TFLite's float hard-swish reference produces; special-casing -Inf would put a per-element select in the hot loop for an input no finite model produces. The scalar and MVE legs agree on the NaN-ness and on +/-Inf; NaN payload bits may differ between legs.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input samples. |
| output | float32_t * | out | Pointer to the output samples. |
| size | int32_t | in | Number of elements to process. Must be at least 1. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:680`
## arm_nn_activation_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_activation_f16(
const float16_t *input,
float16_t *output,
int32_t size,
arm_nn_activation_type_flt type,
float16_t act_param
)
```
Elementwise activation.
:::note
The RELU, RELU6 and LEAKY_RELU legs classify NaN on the integer bit pattern (#380 / #382), so a NaN input comes back as NaN at every optimization level on the gated toolchains, including the shipped -Ofast. This holds on both the scalar and the MVE (cortex-m55) build paths; the MVE RELU/RELU6 legs restore the NaN lanes that vmaxnmq/vminnmq suppress. The MVE TANH leg returns a NaN input unchanged and keeps the sign of zero, also decided on the bit pattern (#635). The scalar TANH leg returns NaN where there is no hardware floating point (__ARM_FP undefined, e.g. Cortex-M0), where it too classifies NaN on the bit pattern (quieting a signalling NaN), and elsewhere only in builds without -ffinite-math-only. SIGMOID and HARDSWISH are outside this contract; see the per-helper notes in `Include/Internal/arm_nn_activation_flt.h`.
:::
:::note
The HARDSWISH leg's scalar helper (arm_nn_hardswish_scalar_f32, serving every build that does not take the MVE float path no MVE float support, or MVE present but not used, e.g. under ARM_MATH_AUTOVECTORIZE) keeps the legacy separately rounded multiply-and-add gate and can differ by an ulp in the curved region from the standalone `arm_hard_swish_f32`, whose gate is a correctly rounded fma; the mux's MVE helper (arm_nn_vhardswish_mve_f32) uses vfmaq and agrees with that kernel. Callers that need bit-exact, leg-agreeing hard swish or the documented NaN/Inf contract should call `arm_hard_swish_f32` directly.
:::
:::note
The RELU, RELU6 and LEAKY_RELU legs classify NaN on the integer bit pattern (#380 / #382), so a NaN input comes back as NaN at every optimization level on the gated toolchains, including the shipped -Ofast. This holds uniformly across build paths: the scalar path serves every build without MVE float16 (and LEAKY_RELU on MVE builds too), while the MVE RELU/RELU6 legs (cortex-m55) restore the NaN lanes that vmaxnmq/vminnmq suppress, using the same integer-domain lane classification as the elementwise clamps. SIGMOID and HARDSWISH are outside this contract; see the per-helper notes in `Include/Internal/arm_nn_activation_flt.h`.
:::
:::note
TANH propagates NaN on both scalar and MVE paths, preserves the sign of zero, and maps +/-Inf to +/-1, including under -Ofast. NaN payload, sign and signaling state are not specified. The finite LUT interpolation may round differently across paths; bitwise scalar/MVE agreement is not required. Caller FP control settings are not changed.
:::
:::note
Both legs of the HARDSWISH mux evaluate natively in float16 the scalar helper (arm_nn_hardswish_scalar_f16) with a separately rounded multiply-and-add gate, the MVE helper (arm_nn_vhardswish_mve_f16) with a float16 vfmaq so either can differ by an ulp from the scalar leg of the standalone `arm_hard_swish_f16`, which computes in float32 with an fma gate and rounds to float16 once. Callers that need the documented NaN/Inf contract should call `arm_hard_swish_f16` directly.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input samples. |
| output | float16_t * | out | Pointer to the output samples. |
| size | int32_t | in | Number of elements to process. |
| type | arm_nn_activation_type_flt | in | Activation selector. |
| act_param | float16_t | in | Extra activation parameter. Used for parameterized activations such as leaky ReLU. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2750`
## arm_prelu_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_prelu_f16(
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *alpha_dims,
const float16_t *alpha,
const cmsis_nn_dims *output_dims,
float16_t *output
)
```
Parametric ReLU for float32 data.
Computes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Must equal output_dims. |
| input | const float16_t * | in | Pointer to the input tensor. |
| alpha_dims | const cmsis_nn_dims * | in | Alpha tensor dimensions. |
| alpha | const float16_t * | in | Pointer to the alpha (slope) tensor. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2759`
## arm_hard_swish_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_hard_swish_f16(const float16_t *input, float16_t *output, int32_t size)
```
Hard swish activation for float16 data.
Computes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise. The scalar leg widens each element to float32, evaluates the gate and the product there exactly as in `arm_hard_swish_f32`, and narrows only the final product, so it is single-rounded. The MVE (cortex-m55) leg evaluates the same expression in float16 throughout, scaling the gate by 1/6 before the product so that the multiplier stays in [0, 1]; it rounds the gate and the product separately and so can sit up to 2 float16 ulp away from the scalar leg in the curved region -3 < x < 3. The saturated regions are exact and identical on both legs (x >= 3 returns x bit-exactly, x <= -3 returns zero), as is the NaN/Inf behavior below; NaN lanes agree in NaN-ness but not necessarily in payload. In-place operation (output == input) is supported on both legs.
:::note
NaN and Inf behave as in `arm_hard_swish_f32`, at every optimization level on the gated toolchains (see docs/guides/verification.md): NaN propagates through the final multiply (NaN-ness only, not a particular payload), +Inf returns +Inf, and -Inf returns NaN because the gate is 0 there and (-Inf) * 0 is NaN by IEEE 754, matching TFLite's float hard-swish reference rather than the mathematical limit 0.
:::
:::note
Nothing in this kernel converts between half and single precision any more, and the float16 kernels that still do are not tied to a particular assembler: they go through `Include/Internal/arm_nn_vcvt_f16.h`, which emits the scalar form of VCVTB/VCVTT wherever the vector form would be mis-encoded (binutils below 2.43). Under CMake the probe measures the assembler in use and selects the form; a build that never runs it the CMSIS-Pack `Source` Cvariant, `module.mk`, or a CMake project that wires its architecture flags where the probe cannot read them falls back to the compiler major, which is right for every Arm GNU release and wrong only for a GCC 14 or newer driver paired by hand with an older binutils. Check `as --version` if you assembled that pair yourself. See docs/guides/toolchains.md.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input samples. |
| output | float16_t * | out | Pointer to the output samples. |
| size | int32_t | in | Number of elements to process. Must be at least 1. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2803`
---
# heliaCORE.Comparison
Comparison operators with optional broadcasting support.
---
# heliaCORE.FC
Collection of fully-connected and matrix multiplication functions.
Fully-connected layer is basically a matrix-vector multiplication with bias. The matrix is the weights and the input/output vectors are the activation values. Supported {weight, activation} precisions include {8-bit, 8-bit} and {8-bit, 16-bit}
## arm_fully_connected_nhwc_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_nhwc_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params_f32 *fc_params,
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *filter_dims,
const float32_t *kernel,
const cmsis_nn_dims *bias_dims,
const float32_t *bias,
const cmsis_nn_dims *output_dims,
float32_t *output
)
```
Fully connected layer, NHWC layout.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f32 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1056`
## arm_fully_connected_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params_f32 *fc_params,
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *filter_dims,
const float32_t *kernel,
const cmsis_nn_dims *bias_dims,
const float32_t *bias,
const cmsis_nn_dims *output_dims,
float32_t *output,
arm_nn_tensor_layout layout
)
```
Fully connected layer, dispatch by layout.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f32 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1084`
## arm_fully_connected_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_fully_connected_f32_get_buffer_size(
const cmsis_nn_fc_params_f32 *fc_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the temporary buffer size required by the fully connected layer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| fc_params | const cmsis_nn_fc_params_f32 * | in | Fully connected parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:1107`
## arm_batch_matmul_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_batch_matmul_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_bmm_params_f32 *bmm_params,
const cmsis_nn_dims *input_lhs_dims,
const float32_t *input_lhs,
const cmsis_nn_dims *input_rhs_dims,
const float32_t *input_rhs,
const cmsis_nn_dims *output_dims,
float32_t *output
)
```
Batched matrix multiplication.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| bmm_params | const cmsis_nn_bmm_params_f32 * | in | Batch matmul parameters and activation clamp. |
| input_lhs_dims | const cmsis_nn_dims * | in | Left-hand-side input tensor dimensions. |
| input_lhs | const float32_t * | in | Pointer to the left-hand-side input tensor. |
| input_rhs_dims | const cmsis_nn_dims * | in | Right-hand-side input tensor dimensions. |
| input_rhs | const float32_t * | in | Pointer to the right-hand-side input tensor. With `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` each `[K, N]` matrix occupies `K * ceil(N / block) * block` elements (block is 4 for float32, 8 for float16) and consecutive batch matrices are stored back to back at that stride. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1491`
## arm_batch_matmul_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_batch_matmul_f32_get_buffer_size(
const cmsis_nn_bmm_params_f32 *bmm_params,
const cmsis_nn_dims *input_lhs_dims,
const cmsis_nn_dims *input_rhs_dims,
const cmsis_nn_dims *output_dims
)
```
Get the temporary buffer size required by batched matrix multiplication.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bmm_params | const cmsis_nn_bmm_params_f32 * | in | Batch matmul parameters. |
| input_lhs_dims | const cmsis_nn_dims * | in | Left-hand-side input tensor dimensions. |
| input_rhs_dims | const cmsis_nn_dims * | in | Right-hand-side input tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:1510`
## arm_fully_connected_nhwc_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_nhwc_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params_f16 *fc_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output
)
```
Fully connected layer, NHWC layout.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3067`
## arm_fully_connected_nhwc_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_nhwc_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params_f16 *fc_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output
)
```
Fully connected layer, NHWC layout.
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_fully_connected_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_fully_connected_nhwc_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3086`
## arm_fully_connected_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params_f16 *fc_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output,
arm_nn_tensor_layout layout
)
```
Fully connected layer, dispatch by layout.
:::note
Accumulation width follows the matmul helper the weight format selects (arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16): the scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457); the MVE legs use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3110`
## arm_fully_connected_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params_f16 *fc_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output,
arm_nn_tensor_layout layout
)
```
Fully connected layer, dispatch by layout.
:::note
Accumulation width follows the matmul helper the weight format selects (arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16): the scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457); the MVE legs use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout.
:::
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_fully_connected_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_fully_connected_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3130`
## arm_fully_connected_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_fully_connected_f16_get_buffer_size(
const cmsis_nn_fc_params_f16 *fc_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the temporary buffer size required by the fully connected layer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:3145`
## arm_batch_matmul_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_batch_matmul_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_bmm_params_f16 *bmm_params,
const cmsis_nn_dims *input_lhs_dims,
const float16_t *input_lhs,
const cmsis_nn_dims *input_rhs_dims,
const float16_t *input_rhs,
const cmsis_nn_dims *output_dims,
float16_t *output
)
```
Batched matrix multiplication.
:::note
Accumulation width. Without adjoints the product goes through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16 and so takes their rule: on the MVE legs a reduction of more than 32 taps per output accumulates blockwise (AmbiqAI/ns-cmsis-nn#586), in float16 up to 32. The adjoint paths accumulate in float16 throughout. There is no `_acc16` entry; a caller that needs float16 lanes on a long reduction calls arm_nn_mat_mult_nt_t_f16_acc16 / arm_nn_mat_mult_nt_n_packed_f16_acc16 per batch.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| bmm_params | const cmsis_nn_bmm_params_f16 * | in | Batch matmul parameters and activation clamp. |
| input_lhs_dims | const cmsis_nn_dims * | in | Left-hand-side input tensor dimensions. |
| input_lhs | const float16_t * | in | Pointer to the left-hand-side input tensor. |
| input_rhs_dims | const cmsis_nn_dims * | in | Right-hand-side input tensor dimensions. |
| input_rhs | const float16_t * | in | Pointer to the right-hand-side input tensor. With `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` each `[K, N]` matrix occupies `K * ceil(N / block) * block` elements (block is 4 for float32, 8 for float16) and consecutive batch matrices are stored back to back at that stride. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3327`
## arm_batch_matmul_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_batch_matmul_f16_get_buffer_size(
const cmsis_nn_bmm_params_f16 *bmm_params,
const cmsis_nn_dims *input_lhs_dims,
const cmsis_nn_dims *input_rhs_dims,
const cmsis_nn_dims *output_dims
)
```
Get the temporary buffer size required by batched matrix multiplication.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bmm_params | const cmsis_nn_bmm_params_f16 * | in | Batch matmul parameters. |
| input_lhs_dims | const cmsis_nn_dims * | in | Left-hand-side input tensor dimensions. |
| input_rhs_dims | const cmsis_nn_dims * | in | Right-hand-side input tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:3339`
---
# heliaCORE.Gather
## arm_gather_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_gather_f32(
const float32_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_params *params,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Gather contiguous slices along an axis.
Data rank is 1..4 and indices rank is 0..4; rank-0 indices contain one index. Negative axis normalizes by input_rank; negative batch_dims normalizes by coords_rank. After normalization, 0 <= batch_dims <= coords_rank and batch_dims <= axis < input_rank. Leading batch dimensions must match. The inferred output shape is input_shape[:axis] + indices_shape[batch_dims:] + input_shape[axis + 1:].
Shapes use the first rank fields of `cmsis_nn_dims` in n, h, w, c order; unused fields are ignored. The inferred output rank must be 0..4, and output_dims must match its leading dimensions. A rank-0 output is one element. All dimension extents must be nonnegative. Input, index and output buffer byte counts must each fit INT32_MAX; this is a CORE capacity limit.
All metadata pointers are required. A NULL data, indices or output pointer is accepted only when that respective buffer has zero elements. Valid empty calls copy nothing. All supplied indices are checked, even for empty output. Coordinates must be nonnegative and below their corresponding axis extent. Invalid metadata or indices return ARG_ERROR without changing output.
This operation preserves all bits, including NaN payloads, signed zero and subnormals, independently of floating-point controls. Buffers must not overlap. No scratch buffer is required. Portable copies use existing MVE copy paths when enabled.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Input data buffer. |
| input_dims | const cmsis_nn_dims * | in | Input shape in leading-dimension order. |
| indices_data | const int32_t * | in | Signed 32-bit indices. |
| indices_dims | const cmsis_nn_dims * | in | Indices shape in leading-dimension order. |
| params | const cmsis_nn_gather_params * | in | Ranks and gathering parameters. |
| output_data | float32_t * | out | Output data buffer. |
| output_dims | const cmsis_nn_dims * | in | Inferred output shape in leading-dimension order. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:2135`
## arm_gather_nd_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_gather_nd_f32(
const float32_t *params_data,
const cmsis_nn_dims *params_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_nd_params *params,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Gather contiguous slices using coordinate tuples.
Data rank is 1..4 and indices rank is 1..4. The final indices dimension is the tuple width, which must be at least one. batch_dims is a TensorFlow-style extension (not a LiteRT builtin option): 0 <= batch_dims < indices_rank, batch_dims < params_rank, and batch_dims + tuple_width <= params_rank. Leading batch dimensions must match. The inferred output shape is indices_shape[:-1] + params_shape[batch_dims + tuple_width:]. Empty data with a nonempty index buffer is rejected.
Shapes use the first rank fields of `cmsis_nn_dims` in n, h, w, c order; unused fields are ignored. The inferred output rank must be 0..4, and output_dims must match its leading dimensions. A rank-0 output is one element. All dimension extents must be nonnegative. Input, index and output buffer byte counts must each fit INT32_MAX; this is a CORE capacity limit.
All metadata pointers are required. A NULL data, indices or output pointer is accepted only when that respective buffer has zero elements. Valid empty calls copy nothing. All supplied indices are checked, even for empty output. Coordinates must be nonnegative and below their corresponding axis extent. Invalid metadata or indices return ARG_ERROR without changing output.
This operation preserves all bits, including NaN payloads, signed zero and subnormals, independently of floating-point controls. Buffers must not overlap. No scratch buffer is required. Portable copies use existing MVE copy paths when enabled.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| params_data | const float32_t * | in | Input data buffer. |
| params_dims | const cmsis_nn_dims * | in | Input shape in leading-dimension order. |
| indices_data | const int32_t * | in | Signed 32-bit indices. |
| indices_dims | const cmsis_nn_dims * | in | Indices shape in leading-dimension order. |
| params | const cmsis_nn_gather_nd_params * | in | Ranks and gathering parameters. |
| output_data | float32_t * | out | Output data buffer. |
| output_dims | const cmsis_nn_dims * | in | Inferred output shape in leading-dimension order. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:2180`
## arm_gather_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_gather_f16(
const float16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_params *params,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Gather contiguous slices along an axis.
Data rank is 1..4 and indices rank is 0..4; rank-0 indices contain one index. Negative axis normalizes by input_rank; negative batch_dims normalizes by coords_rank. After normalization, 0 <= batch_dims <= coords_rank and batch_dims <= axis < input_rank. Leading batch dimensions must match. The inferred output shape is input_shape[:axis] + indices_shape[batch_dims:] + input_shape[axis + 1:].
Shapes use the first rank fields of `cmsis_nn_dims` in n, h, w, c order; unused fields are ignored. The inferred output rank must be 0..4, and output_dims must match its leading dimensions. A rank-0 output is one element. All dimension extents must be nonnegative. Input, index and output buffer byte counts must each fit INT32_MAX; this is a CORE capacity limit.
All metadata pointers are required. A NULL data, indices or output pointer is accepted only when that respective buffer has zero elements. Valid empty calls copy nothing. All supplied indices are checked, even for empty output. Coordinates must be nonnegative and below their corresponding axis extent. Invalid metadata or indices return ARG_ERROR without changing output.
This operation preserves all bits, including NaN payloads, signed zero and subnormals, independently of floating-point controls. Buffers must not overlap. No scratch buffer is required. Portable copies use existing MVE copy paths when enabled.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Input data buffer. |
| input_dims | const cmsis_nn_dims * | in | Input shape in leading-dimension order. |
| indices_data | const int32_t * | in | Signed 32-bit indices. |
| indices_dims | const cmsis_nn_dims * | in | Indices shape in leading-dimension order. |
| params | const cmsis_nn_gather_params * | in | Ranks and gathering parameters. |
| output_data | float16_t * | out | Output data buffer. |
| output_dims | const cmsis_nn_dims * | in | Inferred output shape in leading-dimension order. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:3866`
## arm_gather_nd_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_gather_nd_f16(
const float16_t *params_data,
const cmsis_nn_dims *params_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_nd_params *params,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Gather contiguous slices using coordinate tuples.
Data rank is 1..4 and indices rank is 1..4. The final indices dimension is the tuple width, which must be at least one. batch_dims is a TensorFlow-style extension (not a LiteRT builtin option): 0 <= batch_dims < indices_rank, batch_dims < params_rank, and batch_dims + tuple_width <= params_rank. Leading batch dimensions must match. The inferred output shape is indices_shape[:-1] + params_shape[batch_dims + tuple_width:]. Empty data with a nonempty index buffer is rejected.
Shapes use the first rank fields of `cmsis_nn_dims` in n, h, w, c order; unused fields are ignored. The inferred output rank must be 0..4, and output_dims must match its leading dimensions. A rank-0 output is one element. All dimension extents must be nonnegative. Input, index and output buffer byte counts must each fit INT32_MAX; this is a CORE capacity limit.
All metadata pointers are required. A NULL data, indices or output pointer is accepted only when that respective buffer has zero elements. Valid empty calls copy nothing. All supplied indices are checked, even for empty output. Coordinates must be nonnegative and below their corresponding axis extent. Invalid metadata or indices return ARG_ERROR without changing output.
This operation preserves all bits, including NaN payloads, signed zero and subnormals, independently of floating-point controls. Buffers must not overlap. No scratch buffer is required. Portable copies use existing MVE copy paths when enabled.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| params_data | const float16_t * | in | Input data buffer. |
| params_dims | const cmsis_nn_dims * | in | Input shape in leading-dimension order. |
| indices_data | const int32_t * | in | Signed 32-bit indices. |
| indices_dims | const cmsis_nn_dims * | in | Indices shape in leading-dimension order. |
| params | const cmsis_nn_gather_nd_params * | in | Ranks and gathering parameters. |
| output_data | float16_t * | out | Output data buffer. |
| output_dims | const cmsis_nn_dims * | in | Inferred output shape in leading-dimension order. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:3911`
---
# heliaCORE.genPrivTypes
Data structure types used by private functions.
## arm_nnword
`struct` · `c`
```c
union arm_nnword
```
Union for SIMD access of q31/s16/s8 types.
Source: `Include/arm_nnsupportfunctions.h:596`
### arm_nnword::word
`attribute` · `c`
```c
int32_t word
```
q31 type
Source: `Include/arm_nnsupportfunctions.h:598`
### arm_nnword::half_words
`attribute` · `c`
```c
int16_t half_words[2]
```
s16 type
Source: `Include/arm_nnsupportfunctions.h:600`
### arm_nnword::bytes
`attribute` · `c`
```c
int8_t bytes[4]
```
s8 type
Source: `Include/arm_nnsupportfunctions.h:602`
## arm_nn_double
`struct` · `c`
```c
struct arm_nn_double
```
Union for data type long long.
Source: `Include/arm_nnsupportfunctions.h:609`
### arm_nn_double::low
`attribute` · `c`
```c
uint32_t low
```
Source: `Include/arm_nnsupportfunctions.h:611`
### arm_nn_double::high
`attribute` · `c`
```c
int32_t high
```
Source: `Include/arm_nnsupportfunctions.h:612`
## arm_nn_long_long
`struct` · `c`
```c
union arm_nn_long_long
```
Source: `Include/arm_nnsupportfunctions.h:615`
### arm_nn_long_long::long_long
`attribute` · `c`
```c
int64_t long_long
```
Source: `Include/arm_nnsupportfunctions.h:617`
### arm_nn_long_long::word
`attribute` · `c`
```c
struct arm_nn_double word
```
Source: `Include/arm_nnsupportfunctions.h:618`
---
# heliaCORE.genPubTypes
Enums and Data Structures used in public API.
## cmsis_nn_tile
`struct` · `c`
```c
struct cmsis_nn_tile
```
CMSIS-NN object to contain the width and height of a tile
Source: `Include/arm_nn_types.h:105`
### cmsis_nn_tile::w
`attribute` · `c`
```c
int32_t w
```
Width
Source: `Include/arm_nn_types.h:107`
### cmsis_nn_tile::h
`attribute` · `c`
```c
int32_t h
```
Height
Source: `Include/arm_nn_types.h:108`
## cmsis_nn_context
`struct` · `c`
```c
struct cmsis_nn_context
```
CMSIS-NN object used for the function context.
Source: `Include/arm_nn_types.h:112`
### cmsis_nn_context::buf
`attribute` · `c`
```c
void * buf
```
Pointer to a buffer needed for the optimization
Source: `Include/arm_nn_types.h:114`
### cmsis_nn_context::size
`attribute` · `c`
```c
int32_t size
```
Buffer size
Source: `Include/arm_nn_types.h:115`
## cmsis_nn_bias_data
`struct` · `c`
```c
struct cmsis_nn_bias_data
```
CMSIS-NN object used to hold bias data for int16 variants.
Source: `Include/arm_nn_types.h:119`
### cmsis_nn_bias_data::data
`attribute` · `c`
```c
const void * data
```
Pointer to bias data
Source: `Include/arm_nn_types.h:121`
### cmsis_nn_bias_data::is_int32_bias
`attribute` · `c`
```c
const bool is_int32_bias
```
Indicate type of bias data. True means int32 else int64
Source: `Include/arm_nn_types.h:122`
## cmsis_nn_dims
`struct` · `c`
```c
struct cmsis_nn_dims
```
CMSIS-NN object to contain the dimensions of the tensors
Source: `Include/arm_nn_types.h:126`
### cmsis_nn_dims::n
`attribute` · `c`
```c
int32_t n
```
Generic dimension to contain either the batch size or output channels. Please refer to the function documentation for more information
Source: `Include/arm_nn_types.h:128`
### cmsis_nn_dims::h
`attribute` · `c`
```c
int32_t h
```
Height
Source: `Include/arm_nn_types.h:130`
### cmsis_nn_dims::w
`attribute` · `c`
```c
int32_t w
```
Width
Source: `Include/arm_nn_types.h:131`
### cmsis_nn_dims::c
`attribute` · `c`
```c
int32_t c
```
Input channels
Source: `Include/arm_nn_types.h:132`
## cmsis_nn_lstm_dims
`struct` · `c`
```c
struct cmsis_nn_lstm_dims
```
CMSIS-NN object to contain LSTM specific input parameters related to dimensions
Source: `Include/arm_nn_types.h:136`
### cmsis_nn_lstm_dims::max_time
`attribute` · `c`
```c
int32_t max_time
```
Source: `Include/arm_nn_types.h:138`
### cmsis_nn_lstm_dims::num_inputs
`attribute` · `c`
```c
int32_t num_inputs
```
Source: `Include/arm_nn_types.h:139`
### cmsis_nn_lstm_dims::num_batches
`attribute` · `c`
```c
int32_t num_batches
```
Source: `Include/arm_nn_types.h:140`
### cmsis_nn_lstm_dims::num_outputs
`attribute` · `c`
```c
int32_t num_outputs
```
Source: `Include/arm_nn_types.h:141`
## cmsis_nn_per_channel_quant_params
`struct` · `c`
```c
struct cmsis_nn_per_channel_quant_params
```
CMSIS-NN object for the per-channel quantization parameters
Source: `Include/arm_nn_types.h:145`
### cmsis_nn_per_channel_quant_params::multiplier
`attribute` · `c`
```c
int32_t * multiplier
```
Multiplier values
Source: `Include/arm_nn_types.h:147`
### cmsis_nn_per_channel_quant_params::shift
`attribute` · `c`
```c
int32_t * shift
```
Shift values
Source: `Include/arm_nn_types.h:148`
## cmsis_nn_per_tensor_quant_params
`struct` · `c`
```c
struct cmsis_nn_per_tensor_quant_params
```
CMSIS-NN object for the per-tensor quantization parameters
Source: `Include/arm_nn_types.h:152`
### cmsis_nn_per_tensor_quant_params::multiplier
`attribute` · `c`
```c
int32_t multiplier
```
Multiplier value
Source: `Include/arm_nn_types.h:154`
### cmsis_nn_per_tensor_quant_params::shift
`attribute` · `c`
```c
int32_t shift
```
Shift value
Source: `Include/arm_nn_types.h:155`
## cmsis_nn_quant_params
`struct` · `c`
```c
struct cmsis_nn_quant_params
```
CMSIS-NN object for quantization parameters. This struct supports both per-tensor and per-channels requantization and is recommended for new operators.
Source: `Include/arm_nn_types.h:162`
### cmsis_nn_quant_params::multiplier
`attribute` · `c`
```c
int32_t * multiplier
```
Multiplier values
Source: `Include/arm_nn_types.h:164`
### cmsis_nn_quant_params::shift
`attribute` · `c`
```c
int32_t * shift
```
Shift values
Source: `Include/arm_nn_types.h:165`
### cmsis_nn_quant_params::is_per_channel
`attribute` · `c`
```c
int32_t is_per_channel
```
Source: `Include/arm_nn_types.h:166`
## cmsis_nn_activation
`struct` · `c`
```c
struct cmsis_nn_activation
```
CMSIS-NN object for the quantized Relu activation
Source: `Include/arm_nn_types.h:170`
### cmsis_nn_activation::min
`attribute` · `c`
```c
int32_t min
```
Min value used to clamp the result
Source: `Include/arm_nn_types.h:172`
### cmsis_nn_activation::max
`attribute` · `c`
```c
int32_t max
```
Max value used to clamp the result
Source: `Include/arm_nn_types.h:173`
## cmsis_nn_conv_params
`struct` · `c`
```c
struct cmsis_nn_conv_params
```
CMSIS-NN object for the convolution layer parameters
Source: `Include/arm_nn_types.h:177`
### cmsis_nn_conv_params::input_offset
`attribute` · `c`
```c
int32_t input_offset
```
The negative of the zero value for the input tensor
Source: `Include/arm_nn_types.h:179`
### cmsis_nn_conv_params::output_offset
`attribute` · `c`
```c
int32_t output_offset
```
The negative of the zero value for the output tensor
Source: `Include/arm_nn_types.h:180`
### cmsis_nn_conv_params::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Source: `Include/arm_nn_types.h:181`
### cmsis_nn_conv_params::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Source: `Include/arm_nn_types.h:182`
### cmsis_nn_conv_params::dilation
`attribute` · `c`
```c
cmsis_nn_tile dilation
```
Source: `Include/arm_nn_types.h:183`
### cmsis_nn_conv_params::activation
`attribute` · `c`
```c
cmsis_nn_activation activation
```
Source: `Include/arm_nn_types.h:184`
## cmsis_nn_transpose_conv_params
`struct` · `c`
```c
struct cmsis_nn_transpose_conv_params
```
CMSIS-NN object for the transpose convolution layer parameters
Source: `Include/arm_nn_types.h:188`
### cmsis_nn_transpose_conv_params::input_offset
`attribute` · `c`
```c
int32_t input_offset
```
The negative of the zero value for the input tensor
Source: `Include/arm_nn_types.h:190`
### cmsis_nn_transpose_conv_params::output_offset
`attribute` · `c`
```c
int32_t output_offset
```
The negative of the zero value for the output tensor
Source: `Include/arm_nn_types.h:191`
### cmsis_nn_transpose_conv_params::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Source: `Include/arm_nn_types.h:192`
### cmsis_nn_transpose_conv_params::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Source: `Include/arm_nn_types.h:193`
### cmsis_nn_transpose_conv_params::padding_offsets
`attribute` · `c`
```c
cmsis_nn_tile padding_offsets
```
Source: `Include/arm_nn_types.h:194`
### cmsis_nn_transpose_conv_params::dilation
`attribute` · `c`
```c
cmsis_nn_tile dilation
```
Source: `Include/arm_nn_types.h:195`
### cmsis_nn_transpose_conv_params::activation
`attribute` · `c`
```c
cmsis_nn_activation activation
```
Source: `Include/arm_nn_types.h:196`
## cmsis_nn_dw_conv_params
`struct` · `c`
```c
struct cmsis_nn_dw_conv_params
```
CMSIS-NN object for the depthwise convolution layer parameters
Source: `Include/arm_nn_types.h:200`
### cmsis_nn_dw_conv_params::input_offset
`attribute` · `c`
```c
int32_t input_offset
```
The negative of the zero value for the input tensor
Source: `Include/arm_nn_types.h:202`
### cmsis_nn_dw_conv_params::output_offset
`attribute` · `c`
```c
int32_t output_offset
```
The negative of the zero value for the output tensor
Source: `Include/arm_nn_types.h:203`
### cmsis_nn_dw_conv_params::ch_mult
`attribute` · `c`
```c
int32_t ch_mult
```
Channel Multiplier. ch_mult * in_ch = out_ch
Source: `Include/arm_nn_types.h:204`
### cmsis_nn_dw_conv_params::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Source: `Include/arm_nn_types.h:205`
### cmsis_nn_dw_conv_params::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Source: `Include/arm_nn_types.h:206`
### cmsis_nn_dw_conv_params::dilation
`attribute` · `c`
```c
cmsis_nn_tile dilation
```
Source: `Include/arm_nn_types.h:207`
### cmsis_nn_dw_conv_params::activation
`attribute` · `c`
```c
cmsis_nn_activation activation
```
Source: `Include/arm_nn_types.h:208`
## cmsis_nn_pool_params
`struct` · `c`
```c
struct cmsis_nn_pool_params
```
CMSIS-NN object for pooling layer parameters
Source: `Include/arm_nn_types.h:212`
### cmsis_nn_pool_params::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Source: `Include/arm_nn_types.h:214`
### cmsis_nn_pool_params::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Source: `Include/arm_nn_types.h:215`
### cmsis_nn_pool_params::activation
`attribute` · `c`
```c
cmsis_nn_activation activation
```
Source: `Include/arm_nn_types.h:216`
## cmsis_nn_gather_params
`struct` · `c`
```c
struct cmsis_nn_gather_params
```
CMSIS-NN object for the gather operator
Source: `Include/arm_nn_types.h:220`
### cmsis_nn_gather_params::axis
`attribute` · `c`
```c
int32_t axis
```
Axis to gather from. Supports negative indexing.
Source: `Include/arm_nn_types.h:222`
### cmsis_nn_gather_params::batch_dims
`attribute` · `c`
```c
int32_t batch_dims
```
Number of leading batch dimensions
Source: `Include/arm_nn_types.h:223`
### cmsis_nn_gather_params::input_rank
`attribute` · `c`
```c
int32_t input_rank
```
Rank of the input tensor (range: [1, 4])
Source: `Include/arm_nn_types.h:224`
### cmsis_nn_gather_params::coords_rank
`attribute` · `c`
```c
int32_t coords_rank
```
Rank of the coordinate tensor ([1, 4]; float gather also accepts scalar rank 0)
Source: `Include/arm_nn_types.h:225`
## cmsis_nn_gather_nd_params
`struct` · `c`
```c
struct cmsis_nn_gather_nd_params
```
CMSIS-NN object for the gather_nd operator
Source: `Include/arm_nn_types.h:229`
### cmsis_nn_gather_nd_params::params_rank
`attribute` · `c`
```c
int32_t params_rank
```
Rank of the params tensor (range: [1, 4])
Source: `Include/arm_nn_types.h:231`
### cmsis_nn_gather_nd_params::indices_rank
`attribute` · `c`
```c
int32_t indices_rank
```
Rank of the indices tensor (range: [1, 4])
Source: `Include/arm_nn_types.h:232`
### cmsis_nn_gather_nd_params::batch_dims
`attribute` · `c`
```c
int32_t batch_dims
```
Number of batch dimensions
Source: `Include/arm_nn_types.h:233`
## cmsis_nn_tile_params
`struct` · `c`
```c
struct cmsis_nn_tile_params
```
CMSIS-NN object for the tile operator
Source: `Include/arm_nn_types.h:237`
### cmsis_nn_tile_params::rank
`attribute` · `c`
```c
int32_t rank
```
Rank of the input tensor (range: [1, 8])
Source: `Include/arm_nn_types.h:239`
### cmsis_nn_tile_params::input_shape
`attribute` · `c`
```c
const int32_t * input_shape
```
Input shape array (length = rank)
Source: `Include/arm_nn_types.h:240`
### cmsis_nn_tile_params::multiples
`attribute` · `c`
```c
const int32_t * multiples
```
Multiples array (length = rank)
Source: `Include/arm_nn_types.h:241`
## cmsis_nn_broadcast_to_params
`struct` · `c`
```c
struct cmsis_nn_broadcast_to_params
```
CMSIS-NN object for the broadcast_to operator
Source: `Include/arm_nn_types.h:245`
### cmsis_nn_broadcast_to_params::rank
`attribute` · `c`
```c
int32_t rank
```
Rank of input/output tensors (range: [1, 8])
Source: `Include/arm_nn_types.h:247`
### cmsis_nn_broadcast_to_params::input_shape
`attribute` · `c`
```c
const int32_t * input_shape
```
Input shape array (length = rank)
Source: `Include/arm_nn_types.h:248`
### cmsis_nn_broadcast_to_params::output_shape
`attribute` · `c`
```c
const int32_t * output_shape
```
Output (broadcast target) shape array (length = rank)
Source: `Include/arm_nn_types.h:249`
## cmsis_nn_scatter_nd_params
`struct` · `c`
```c
struct cmsis_nn_scatter_nd_params
```
CMSIS-NN object for the scatter_nd operator
Source: `Include/arm_nn_types.h:253`
### cmsis_nn_scatter_nd_params::num_updates
`attribute` · `c`
```c
int32_t num_updates
```
Number of update slices
Source: `Include/arm_nn_types.h:255`
### cmsis_nn_scatter_nd_params::index_depth
`attribute` · `c`
```c
int32_t index_depth
```
Depth of each index vector
Source: `Include/arm_nn_types.h:256`
### cmsis_nn_scatter_nd_params::slice_size
`attribute` · `c`
```c
int32_t slice_size
```
Size of each update slice
Source: `Include/arm_nn_types.h:257`
### cmsis_nn_scatter_nd_params::output_size
`attribute` · `c`
```c
int32_t output_size
```
Total number of elements in output
Source: `Include/arm_nn_types.h:258`
### cmsis_nn_scatter_nd_params::output_strides
`attribute` · `c`
```c
const int32_t * output_strides
```
Strides of the output tensor (length = index_depth)
Source: `Include/arm_nn_types.h:259`
## cmsis_nn_mirror_pad_params
`struct` · `c`
```c
struct cmsis_nn_mirror_pad_params
```
CMSIS-NN object for the mirror_pad operator
Source: `Include/arm_nn_types.h:263`
### cmsis_nn_mirror_pad_params::rank
`attribute` · `c`
```c
int32_t rank
```
Rank of the input tensor (range: [1, 8])
Source: `Include/arm_nn_types.h:265`
### cmsis_nn_mirror_pad_params::input_shape
`attribute` · `c`
```c
const int32_t * input_shape
```
Input shape array (length = rank)
Source: `Include/arm_nn_types.h:266`
### cmsis_nn_mirror_pad_params::output_shape
`attribute` · `c`
```c
const int32_t * output_shape
```
Output shape array (length = rank)
Source: `Include/arm_nn_types.h:267`
### cmsis_nn_mirror_pad_params::pad_before
`attribute` · `c`
```c
const int32_t * pad_before
```
Padding before each dimension (length = rank)
Source: `Include/arm_nn_types.h:268`
### cmsis_nn_mirror_pad_params::mode
`attribute` · `c`
```c
int32_t mode
```
0 = REFLECT, 1 = SYMMETRIC
Source: `Include/arm_nn_types.h:269`
## cmsis_nn_where_params
`struct` · `c`
```c
struct cmsis_nn_where_params
```
CMSIS-NN object for the WHERE operator
Source: `Include/arm_nn_types.h:273`
### cmsis_nn_where_params::rank
`attribute` · `c`
```c
int32_t rank
```
Rank of the condition tensor (range: [1, 8])
Source: `Include/arm_nn_types.h:275`
### cmsis_nn_where_params::shape
`attribute` · `c`
```c
const int32_t * shape
```
Condition tensor shape array (length = rank)
Source: `Include/arm_nn_types.h:276`
## cmsis_nn_select_v2_params
`struct` · `c`
```c
struct cmsis_nn_select_v2_params
```
CMSIS-NN object for the select_v2 operator (with broadcast)
Source: `Include/arm_nn_types.h:280`
### cmsis_nn_select_v2_params::rank
`attribute` · `c`
```c
int32_t rank
```
Rank of the output tensor (range: [1, 8])
Source: `Include/arm_nn_types.h:282`
### cmsis_nn_select_v2_params::output_shape
`attribute` · `c`
```c
const int32_t * output_shape
```
Output shape array (length = rank)
Source: `Include/arm_nn_types.h:283`
### cmsis_nn_select_v2_params::cond_strides
`attribute` · `c`
```c
const int32_t * cond_strides
```
Condition tensor broadcast strides (length = rank)
Source: `Include/arm_nn_types.h:284`
### cmsis_nn_select_v2_params::x_strides
`attribute` · `c`
```c
const int32_t * x_strides
```
X tensor broadcast strides (length = rank)
Source: `Include/arm_nn_types.h:285`
### cmsis_nn_select_v2_params::y_strides
`attribute` · `c`
```c
const int32_t * y_strides
```
Y tensor broadcast strides (length = rank)
Source: `Include/arm_nn_types.h:286`
## cmsis_nn_reverse_sequence_params
`struct` · `c`
```c
struct cmsis_nn_reverse_sequence_params
```
CMSIS-NN object for the reverse_sequence operator
Source: `Include/arm_nn_types.h:290`
### cmsis_nn_reverse_sequence_params::rank
`attribute` · `c`
```c
int32_t rank
```
Rank of the input tensor (range: [1, 8])
Source: `Include/arm_nn_types.h:292`
### cmsis_nn_reverse_sequence_params::shape
`attribute` · `c`
```c
const int32_t * shape
```
Input shape array (length = rank)
Source: `Include/arm_nn_types.h:293`
### cmsis_nn_reverse_sequence_params::seq_dim
`attribute` · `c`
```c
int32_t seq_dim
```
Dimension along which to reverse
Source: `Include/arm_nn_types.h:294`
### cmsis_nn_reverse_sequence_params::batch_dim
`attribute` · `c`
```c
int32_t batch_dim
```
Batch dimension
Source: `Include/arm_nn_types.h:295`
## cmsis_nn_dynamic_update_slice_params
`struct` · `c`
```c
struct cmsis_nn_dynamic_update_slice_params
```
CMSIS-NN object for the dynamic_update_slice operator
Source: `Include/arm_nn_types.h:299`
### cmsis_nn_dynamic_update_slice_params::rank
`attribute` · `c`
```c
int32_t rank
```
Rank of the operand tensor (range: [1, 8])
Source: `Include/arm_nn_types.h:301`
### cmsis_nn_dynamic_update_slice_params::operand_shape
`attribute` · `c`
```c
const int32_t * operand_shape
```
Operand shape array (length = rank)
Source: `Include/arm_nn_types.h:302`
### cmsis_nn_dynamic_update_slice_params::update_shape
`attribute` · `c`
```c
const int32_t * update_shape
```
Update shape array (length = rank)
Source: `Include/arm_nn_types.h:303`
### cmsis_nn_dynamic_update_slice_params::operand_size
`attribute` · `c`
```c
int32_t operand_size
```
Total number of elements in operand
Source: `Include/arm_nn_types.h:304`
### cmsis_nn_dynamic_update_slice_params::update_size
`attribute` · `c`
```c
int32_t update_size
```
Total number of elements in update
Source: `Include/arm_nn_types.h:305`
### cmsis_nn_dynamic_update_slice_params::operand_strides
`attribute` · `c`
```c
const int32_t * operand_strides
```
Strides of the operand tensor (length = rank)
Source: `Include/arm_nn_types.h:306`
## cmsis_nn_fc_params
`struct` · `c`
```c
struct cmsis_nn_fc_params
```
CMSIS-NN object for Fully Connected layer parameters
Source: `Include/arm_nn_types.h:310`
### cmsis_nn_fc_params::input_offset
`attribute` · `c`
```c
int32_t input_offset
```
The negative of the zero value for the input tensor
Source: `Include/arm_nn_types.h:312`
### cmsis_nn_fc_params::filter_offset
`attribute` · `c`
```c
int32_t filter_offset
```
The negative of the zero value for the filter tensor
Source: `Include/arm_nn_types.h:313`
### cmsis_nn_fc_params::output_offset
`attribute` · `c`
```c
int32_t output_offset
```
The negative of the zero value for the output tensor
Source: `Include/arm_nn_types.h:314`
### cmsis_nn_fc_params::activation
`attribute` · `c`
```c
cmsis_nn_activation activation
```
Source: `Include/arm_nn_types.h:315`
## cmsis_nn_bmm_params
`struct` · `c`
```c
struct cmsis_nn_bmm_params
```
CMSIS-NN object for Batch Matmul layer parameters
Source: `Include/arm_nn_types.h:319`
### cmsis_nn_bmm_params::adj_x
`attribute` · `c`
```c
const bool adj_x
```
Source: `Include/arm_nn_types.h:321`
### cmsis_nn_bmm_params::adj_y
`attribute` · `c`
```c
const bool adj_y
```
Source: `Include/arm_nn_types.h:322`
### cmsis_nn_bmm_params::fc_params
`attribute` · `c`
```c
cmsis_nn_fc_params fc_params
```
Source: `Include/arm_nn_types.h:323`
## cmsis_nn_transpose_params
`struct` · `c`
```c
struct cmsis_nn_transpose_params
```
CMSIS-NN object for Transpose layer parameters
Source: `Include/arm_nn_types.h:327`
### cmsis_nn_transpose_params::num_dims
`attribute` · `c`
```c
const int32_t num_dims
```
Source: `Include/arm_nn_types.h:329`
### cmsis_nn_transpose_params::permutations
`attribute` · `c`
```c
const uint32_t * permutations
```
The dimensions applied to the input dimensions
Source: `Include/arm_nn_types.h:330`
## cmsis_nn_resize_params
`struct` · `c`
```c
struct cmsis_nn_resize_params
```
CMSIS-NN object for Resize Nearest Neighbor layer parameters
Source: `Include/arm_nn_types.h:334`
### cmsis_nn_resize_params::align_corners
`attribute` · `c`
```c
bool align_corners
```
Align corners when calculating interpolation
Source: `Include/arm_nn_types.h:336`
### cmsis_nn_resize_params::half_pixel_centers
`attribute` · `c`
```c
bool half_pixel_centers
```
Use half pixel centers when calculating interpolation
Source: `Include/arm_nn_types.h:337`
## cmsis_nn_svdf_params
`struct` · `c`
```c
struct cmsis_nn_svdf_params
```
CMSIS-NN object for SVDF layer parameters
Source: `Include/arm_nn_types.h:341`
### cmsis_nn_svdf_params::rank
`attribute` · `c`
```c
int32_t rank
```
Source: `Include/arm_nn_types.h:343`
### cmsis_nn_svdf_params::input_offset
`attribute` · `c`
```c
int32_t input_offset
```
The negative of the zero value for the input tensor
Source: `Include/arm_nn_types.h:344`
### cmsis_nn_svdf_params::output_offset
`attribute` · `c`
```c
int32_t output_offset
```
The negative of the zero value for the output tensor
Source: `Include/arm_nn_types.h:345`
### cmsis_nn_svdf_params::input_activation
`attribute` · `c`
```c
cmsis_nn_activation input_activation
```
Source: `Include/arm_nn_types.h:346`
### cmsis_nn_svdf_params::output_activation
`attribute` · `c`
```c
cmsis_nn_activation output_activation
```
Source: `Include/arm_nn_types.h:347`
## cmsis_nn_softmax_lut_s16
`struct` · `c`
```c
struct cmsis_nn_softmax_lut_s16
```
CMSIS-NN object for Softmax s16 layer parameters
Source: `Include/arm_nn_types.h:351`
### cmsis_nn_softmax_lut_s16::exp_lut
`attribute` · `c`
```c
const int16_t * exp_lut
```
Source: `Include/arm_nn_types.h:353`
### cmsis_nn_softmax_lut_s16::one_by_one_lut
`attribute` · `c`
```c
const int16_t * one_by_one_lut
```
Source: `Include/arm_nn_types.h:354`
## cmsis_nn_concatenation_params
`struct` · `c`
```c
struct cmsis_nn_concatenation_params
```
Source: `Include/arm_nn_types.h:357`
### cmsis_nn_concatenation_params::axis
`attribute` · `c`
```c
const int32_t axis
```
Source: `Include/arm_nn_types.h:359`
## cmsis_nn_scaling
`struct` · `c`
```c
struct cmsis_nn_scaling
```
CMSIS-NN object for quantization parameters
Source: `Include/arm_nn_types.h:363`
### cmsis_nn_scaling::multiplier
`attribute` · `c`
```c
int32_t multiplier
```
Multiplier value
Source: `Include/arm_nn_types.h:365`
### cmsis_nn_scaling::shift
`attribute` · `c`
```c
int32_t shift
```
Shift value
Source: `Include/arm_nn_types.h:366`
## cmsis_nn_lstm_gate
`struct` · `c`
```c
struct cmsis_nn_lstm_gate
```
CMSIS-NN object for LSTM gate parameters
Source: `Include/arm_nn_types.h:370`
### cmsis_nn_lstm_gate::input_multiplier
`attribute` · `c`
```c
int32_t input_multiplier
```
Source: `Include/arm_nn_types.h:372`
### cmsis_nn_lstm_gate::input_shift
`attribute` · `c`
```c
int32_t input_shift
```
Source: `Include/arm_nn_types.h:373`
### cmsis_nn_lstm_gate::input_weights
`attribute` · `c`
```c
const void * input_weights
```
Source: `Include/arm_nn_types.h:374`
### cmsis_nn_lstm_gate::input_effective_bias
`attribute` · `c`
```c
const void * input_effective_bias
```
Bias added with precomputed kernel_sum * lhs_offset
Source: `Include/arm_nn_types.h:375`
### cmsis_nn_lstm_gate::hidden_multiplier
`attribute` · `c`
```c
int32_t hidden_multiplier
```
Source: `Include/arm_nn_types.h:377`
### cmsis_nn_lstm_gate::hidden_shift
`attribute` · `c`
```c
int32_t hidden_shift
```
Source: `Include/arm_nn_types.h:378`
### cmsis_nn_lstm_gate::hidden_weights
`attribute` · `c`
```c
const void * hidden_weights
```
Source: `Include/arm_nn_types.h:379`
### cmsis_nn_lstm_gate::hidden_effective_bias
`attribute` · `c`
```c
const void * hidden_effective_bias
```
Precomputed kernel_sum * lhs_offset
Source: `Include/arm_nn_types.h:380`
### cmsis_nn_lstm_gate::bias
`attribute` · `c`
```c
const void * bias
```
Source: `Include/arm_nn_types.h:382`
### cmsis_nn_lstm_gate::activation_type
`attribute` · `c`
```c
arm_nn_activation_type activation_type
```
Source: `Include/arm_nn_types.h:383`
## cmsis_nn_lstm_params
`struct` · `c`
```c
struct cmsis_nn_lstm_params
```
CMSIS-NN object for LSTM parameters
Source: `Include/arm_nn_types.h:387`
### cmsis_nn_lstm_params::time_major
`attribute` · `c`
```c
int32_t time_major
```
0 if first dimension is batch, else first dimension is time
Source: `Include/arm_nn_types.h:389`
### cmsis_nn_lstm_params::batch_size
`attribute` · `c`
```c
int32_t batch_size
```
Source: `Include/arm_nn_types.h:390`
### cmsis_nn_lstm_params::time_steps
`attribute` · `c`
```c
int32_t time_steps
```
Source: `Include/arm_nn_types.h:391`
### cmsis_nn_lstm_params::input_size
`attribute` · `c`
```c
int32_t input_size
```
Size of new data input into the LSTM cell
Source: `Include/arm_nn_types.h:392`
### cmsis_nn_lstm_params::hidden_size
`attribute` · `c`
```c
int32_t hidden_size
```
Size of output from the LSTM cell, used as output and recursively into the next time step
Source: `Include/arm_nn_types.h:394`
### cmsis_nn_lstm_params::input_offset
`attribute` · `c`
```c
int32_t input_offset
```
Source: `Include/arm_nn_types.h:396`
### cmsis_nn_lstm_params::forget_to_cell_multiplier
`attribute` · `c`
```c
int32_t forget_to_cell_multiplier
```
Source: `Include/arm_nn_types.h:398`
### cmsis_nn_lstm_params::forget_to_cell_shift
`attribute` · `c`
```c
int32_t forget_to_cell_shift
```
Source: `Include/arm_nn_types.h:399`
### cmsis_nn_lstm_params::input_to_cell_multiplier
`attribute` · `c`
```c
int32_t input_to_cell_multiplier
```
Source: `Include/arm_nn_types.h:400`
### cmsis_nn_lstm_params::input_to_cell_shift
`attribute` · `c`
```c
int32_t input_to_cell_shift
```
Source: `Include/arm_nn_types.h:401`
### cmsis_nn_lstm_params::cell_clip
`attribute` · `c`
```c
int32_t cell_clip
```
Min/max value of cell output
Source: `Include/arm_nn_types.h:402`
### cmsis_nn_lstm_params::cell_scale_power
`attribute` · `c`
```c
int32_t cell_scale_power
```
Source: `Include/arm_nn_types.h:403`
### cmsis_nn_lstm_params::output_multiplier
`attribute` · `c`
```c
int32_t output_multiplier
```
Source: `Include/arm_nn_types.h:405`
### cmsis_nn_lstm_params::output_shift
`attribute` · `c`
```c
int32_t output_shift
```
Source: `Include/arm_nn_types.h:406`
### cmsis_nn_lstm_params::output_offset
`attribute` · `c`
```c
int32_t output_offset
```
Source: `Include/arm_nn_types.h:407`
### cmsis_nn_lstm_params::forget_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate forget_gate
```
Source: `Include/arm_nn_types.h:409`
### cmsis_nn_lstm_params::input_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate input_gate
```
Source: `Include/arm_nn_types.h:410`
### cmsis_nn_lstm_params::cell_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate cell_gate
```
Source: `Include/arm_nn_types.h:411`
### cmsis_nn_lstm_params::output_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate output_gate
```
Source: `Include/arm_nn_types.h:412`
## cmsis_nn_lstm_context
`struct` · `c`
```c
struct cmsis_nn_lstm_context
```
CMSIS-NN object for LSTM scratch buffers.
There is no size field and no runtime enforcement: an undersized temp1 or temp2 is written past on every build target, so size them from the queries below, not by transcribing a formula.
Source: `Include/arm_nn_types.h:420`
### cmsis_nn_lstm_context::temp1
`attribute` · `c`
```c
void * temp1
```
Gate-vector scratch (int16_t elements for both the s8 and s16 layers). Sized by `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` / `arm_lstm_unidirectional_s16_temp1_get_buffer_size()`.
Source: `Include/arm_nn_types.h:422`
### cmsis_nn_lstm_context::temp2
`attribute` · `c`
```c
void * temp2
```
Cell-gate and tanh(cell_state) scratch (int16_t elements for both layers). Sized by `arm_lstm_unidirectional_s8_temp2_get_buffer_size()` / `arm_lstm_unidirectional_s16_temp2_get_buffer_size()`.
Source: `Include/arm_nn_types.h:425`
### cmsis_nn_lstm_context::cell_state
`attribute` · `c`
```c
void * cell_state
```
Cell-state buffer, batch_size * hidden_size int16_t elements for both layers.
Source: `Include/arm_nn_types.h:428`
### cmsis_nn_lstm_context::hidden_state
`attribute` · `c`
```c
void * hidden_state
```
Optional in/out hidden state for streaming; NULL selects stateless.
Source: `Include/arm_nn_types.h:429`
## cmsis_nn_activation_f32
`struct` · `c`
```c
struct cmsis_nn_activation_f32
```
Activation clamp range for floating-point operators.
Source: `Include/arm_nn_types_flt.h:133`
### cmsis_nn_activation_f32::min
`attribute` · `c`
```c
float32_t min
```
Minimum value used to clamp the result.
Source: `Include/arm_nn_types_flt.h:135`
### cmsis_nn_activation_f32::max
`attribute` · `c`
```c
float32_t max
```
Maximum value used to clamp the result.
Source: `Include/arm_nn_types_flt.h:136`
## cmsis_nn_conv_params_f32
`struct` · `c`
```c
struct cmsis_nn_conv_params_f32
```
Convolution parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:142`
### cmsis_nn_conv_params_f32::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Spatial stride.
Source: `Include/arm_nn_types_flt.h:144`
### cmsis_nn_conv_params_f32::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Spatial zero-padding.
Source: `Include/arm_nn_types_flt.h:145`
### cmsis_nn_conv_params_f32::dilation
`attribute` · `c`
```c
cmsis_nn_tile dilation
```
Spatial dilation.
Source: `Include/arm_nn_types_flt.h:146`
### cmsis_nn_conv_params_f32::activation
`attribute` · `c`
```c
cmsis_nn_activation_f32 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:147`
### cmsis_nn_conv_params_f32::weight_format
`attribute` · `c`
```c
arm_nn_weight_format_flt weight_format
```
Filter storage format.
Source: `Include/arm_nn_types_flt.h:148`
## cmsis_nn_transpose_conv_params_f32
`struct` · `c`
```c
struct cmsis_nn_transpose_conv_params_f32
```
Transpose convolution parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:154`
### cmsis_nn_transpose_conv_params_f32::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Spatial stride.
Source: `Include/arm_nn_types_flt.h:156`
### cmsis_nn_transpose_conv_params_f32::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Spatial zero-padding.
Source: `Include/arm_nn_types_flt.h:157`
### cmsis_nn_transpose_conv_params_f32::padding_offsets
`attribute` · `c`
```c
cmsis_nn_tile padding_offsets
```
Output padding adjustment for transpose convolution.
Source: `Include/arm_nn_types_flt.h:158`
### cmsis_nn_transpose_conv_params_f32::dilation
`attribute` · `c`
```c
cmsis_nn_tile dilation
```
Spatial dilation.
Source: `Include/arm_nn_types_flt.h:159`
### cmsis_nn_transpose_conv_params_f32::activation
`attribute` · `c`
```c
cmsis_nn_activation_f32 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:160`
## cmsis_nn_dw_conv_params_f32
`struct` · `c`
```c
struct cmsis_nn_dw_conv_params_f32
```
Depthwise convolution parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:166`
### cmsis_nn_dw_conv_params_f32::ch_mult
`attribute` · `c`
```c
int32_t ch_mult
```
Channel multiplier. `ch_mult * in_ch = out_ch`.
Source: `Include/arm_nn_types_flt.h:168`
### cmsis_nn_dw_conv_params_f32::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Spatial stride.
Source: `Include/arm_nn_types_flt.h:169`
### cmsis_nn_dw_conv_params_f32::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Spatial zero-padding.
Source: `Include/arm_nn_types_flt.h:170`
### cmsis_nn_dw_conv_params_f32::dilation
`attribute` · `c`
```c
cmsis_nn_tile dilation
```
Spatial dilation.
Source: `Include/arm_nn_types_flt.h:171`
### cmsis_nn_dw_conv_params_f32::activation
`attribute` · `c`
```c
cmsis_nn_activation_f32 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:172`
## cmsis_nn_pool_params_f32
`struct` · `c`
```c
struct cmsis_nn_pool_params_f32
```
Pooling parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:178`
### cmsis_nn_pool_params_f32::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Spatial stride.
Source: `Include/arm_nn_types_flt.h:180`
### cmsis_nn_pool_params_f32::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Spatial zero-padding.
Source: `Include/arm_nn_types_flt.h:181`
### cmsis_nn_pool_params_f32::activation
`attribute` · `c`
```c
cmsis_nn_activation_f32 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:182`
## cmsis_nn_fc_params_f32
`struct` · `c`
```c
struct cmsis_nn_fc_params_f32
```
Fully connected layer parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:188`
### cmsis_nn_fc_params_f32::activation
`attribute` · `c`
```c
cmsis_nn_activation_f32 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:190`
### cmsis_nn_fc_params_f32::weight_format
`attribute` · `c`
```c
arm_nn_weight_format_flt weight_format
```
Weight storage format.
Source: `Include/arm_nn_types_flt.h:191`
## cmsis_nn_bmm_params_f32
`struct` · `c`
```c
struct cmsis_nn_bmm_params_f32
```
Batched matrix multiplication parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:197`
### cmsis_nn_bmm_params_f32::adj_x
`attribute` · `c`
```c
const bool adj_x
```
True when the left-hand-side operand is stored transposed.
Source: `Include/arm_nn_types_flt.h:199`
### cmsis_nn_bmm_params_f32::adj_y
`attribute` · `c`
```c
const bool adj_y
```
True when the right-hand-side operand is stored transposed.
Source: `Include/arm_nn_types_flt.h:200`
### cmsis_nn_bmm_params_f32::activation
`attribute` · `c`
```c
cmsis_nn_activation_f32 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:201`
### cmsis_nn_bmm_params_f32::rhs_format
`attribute` · `c`
```c
arm_nn_weight_format_flt rhs_format
```
Right-hand-side operand storage format. `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` is currently supported only when `adj_x == false` and `adj_y == false`.
Source: `Include/arm_nn_types_flt.h:202`
## cmsis_nn_ew_params_f32
`struct` · `c`
```c
struct cmsis_nn_ew_params_f32
```
Elementwise operator parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:211`
### cmsis_nn_ew_params_f32::activation
`attribute` · `c`
```c
cmsis_nn_activation_f32 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:213`
## cmsis_nn_transpose_params_f32
`struct` · `c`
```c
struct cmsis_nn_transpose_params_f32
```
Transpose parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:219`
### cmsis_nn_transpose_params_f32::num_dims
`attribute` · `c`
```c
int32_t num_dims
```
Number of active dimensions in the permutation.
Source: `Include/arm_nn_types_flt.h:221`
### cmsis_nn_transpose_params_f32::perm
`attribute` · `c`
```c
int32_t perm[4]
```
Permutation indices.
Source: `Include/arm_nn_types_flt.h:222`
### cmsis_nn_transpose_params_f32::layout
`attribute` · `c`
```c
arm_nn_tensor_layout layout
```
Layout convention used to interpret tensor dimensions.
Source: `Include/arm_nn_types_flt.h:223`
## cmsis_nn_svdf_params_f32
`struct` · `c`
```c
struct cmsis_nn_svdf_params_f32
```
Singular value decomposition filter parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:229`
### cmsis_nn_svdf_params_f32::rank
`attribute` · `c`
```c
int32_t rank
```
SVDF rank.
Source: `Include/arm_nn_types_flt.h:231`
### cmsis_nn_svdf_params_f32::input_activation
`attribute` · `c`
```c
cmsis_nn_activation_f32 input_activation
```
Clamp range applied after the input projection.
Source: `Include/arm_nn_types_flt.h:232`
### cmsis_nn_svdf_params_f32::output_activation
`attribute` · `c`
```c
cmsis_nn_activation_f32 output_activation
```
Clamp range applied to the final output.
Source: `Include/arm_nn_types_flt.h:233`
## cmsis_nn_lstm_gate_f32
`struct` · `c`
```c
struct cmsis_nn_lstm_gate_f32
```
Read-only weights and bias metadata for one float32 LSTM gate.
Source: `Include/arm_nn_types_flt.h:239`
### cmsis_nn_lstm_gate_f32::input_weights
`attribute` · `c`
```c
const float32_t * input_weights
```
Input-to-gate weight matrix.
Source: `Include/arm_nn_types_flt.h:241`
### cmsis_nn_lstm_gate_f32::hidden_weights
`attribute` · `c`
```c
const float32_t * hidden_weights
```
Hidden-state-to-gate weight matrix.
Source: `Include/arm_nn_types_flt.h:242`
### cmsis_nn_lstm_gate_f32::bias
`attribute` · `c`
```c
const float32_t * bias
```
Optional gate bias vector.
Source: `Include/arm_nn_types_flt.h:243`
### cmsis_nn_lstm_gate_f32::activation_type
`attribute` · `c`
```c
arm_nn_activation_type_flt activation_type
```
Gate activation selector.
Source: `Include/arm_nn_types_flt.h:244`
## cmsis_nn_lstm_params_f32
`struct` · `c`
```c
struct cmsis_nn_lstm_params_f32
```
Parameters for a unidirectional float32 LSTM layer.
Source: `Include/arm_nn_types_flt.h:250`
### cmsis_nn_lstm_params_f32::time_major
`attribute` · `c`
```c
int32_t time_major
```
Non-zero when input/output tensors are time-major.
Source: `Include/arm_nn_types_flt.h:252`
### cmsis_nn_lstm_params_f32::batch_size
`attribute` · `c`
```c
int32_t batch_size
```
Batch size processed per invocation.
Source: `Include/arm_nn_types_flt.h:253`
### cmsis_nn_lstm_params_f32::time_steps
`attribute` · `c`
```c
int32_t time_steps
```
Number of time steps processed per invocation.
Source: `Include/arm_nn_types_flt.h:254`
### cmsis_nn_lstm_params_f32::input_size
`attribute` · `c`
```c
int32_t input_size
```
Input feature size per time step.
Source: `Include/arm_nn_types_flt.h:255`
### cmsis_nn_lstm_params_f32::hidden_size
`attribute` · `c`
```c
int32_t hidden_size
```
Hidden-state size.
Source: `Include/arm_nn_types_flt.h:256`
### cmsis_nn_lstm_params_f32::cell_clip
`attribute` · `c`
```c
float32_t cell_clip
```
Optional cell-state clip value.
Source: `Include/arm_nn_types_flt.h:257`
### cmsis_nn_lstm_params_f32::forget_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate_f32 forget_gate
```
Forget gate weights and activation.
Source: `Include/arm_nn_types_flt.h:259`
### cmsis_nn_lstm_params_f32::input_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate_f32 input_gate
```
Input gate weights and activation.
Source: `Include/arm_nn_types_flt.h:260`
### cmsis_nn_lstm_params_f32::cell_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate_f32 cell_gate
```
Cell-update gate weights and activation.
Source: `Include/arm_nn_types_flt.h:261`
### cmsis_nn_lstm_params_f32::output_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate_f32 output_gate
```
Output gate weights and activation.
Source: `Include/arm_nn_types_flt.h:262`
## cmsis_nn_lstm_context_f32
`struct` · `c`
```c
struct cmsis_nn_lstm_context_f32
```
Scratch and mutable state buffers for a float32 LSTM invocation.
Source: `Include/arm_nn_types_flt.h:268`
### cmsis_nn_lstm_context_f32::temp1
`attribute` · `c`
```c
float32_t * temp1
```
Unused by the current implementation and may be NULL. Sized by `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, which reports 0; size from the query rather than hard-coding NULL if the buffer is arena-allocated.
Source: `Include/arm_nn_types_flt.h:270`
### cmsis_nn_lstm_context_f32::temp2
`attribute` · `c`
```c
float32_t * temp2
```
Unused by the current implementation and may be NULL. Sized by `arm_lstm_unidirectional_f32_temp2_get_buffer_size()`.
Source: `Include/arm_nn_types_flt.h:273`
### cmsis_nn_lstm_context_f32::cell_state
`attribute` · `c`
```c
float32_t * cell_state
```
Mutable cell-state buffer (in/out when streaming).
Source: `Include/arm_nn_types_flt.h:275`
### cmsis_nn_lstm_context_f32::hidden_state
`attribute` · `c`
```c
float32_t * hidden_state
```
Optional in/out hidden state for streaming; NULL selects stateless. Streaming is NULL-gated, so zero-initialise the context (e.g. designated initialisers) to keep legacy 3-field callers stateless. Matches the quantized `cmsis_nn_lstm_context` contract.
Source: `Include/arm_nn_types_flt.h:276`
## cmsis_nn_gru_gate_f32
`struct` · `c`
```c
struct cmsis_nn_gru_gate_f32
```
Weights and biases for a single float32 GRU gate.
The activation (sigmoid for update/reset, tanh for candidate) is implied by the gate's role and is not stored here. The reset-after formulation keeps the input-projection bias and the recurrent-projection bias separate, because the reset gate multiplies the recurrent projection (including its bias) after the matmul.
Source: `Include/arm_nn_types_flt.h:291`
### cmsis_nn_gru_gate_f32::input_weights
`attribute` · `c`
```c
const float32_t * input_weights
```
Input-to-gate weight matrix [hidden_size, input_size].
Source: `Include/arm_nn_types_flt.h:293`
### cmsis_nn_gru_gate_f32::hidden_weights
`attribute` · `c`
```c
const float32_t * hidden_weights
```
Hidden-to-gate weight matrix [hidden_size, hidden_size].
Source: `Include/arm_nn_types_flt.h:294`
### cmsis_nn_gru_gate_f32::input_bias
`attribute` · `c`
```c
const float32_t * input_bias
```
Optional input-projection bias [hidden_size]. May be NULL.
Source: `Include/arm_nn_types_flt.h:295`
### cmsis_nn_gru_gate_f32::hidden_bias
`attribute` · `c`
```c
const float32_t * hidden_bias
```
Optional recurrent-projection bias [hidden_size]. May be NULL.
Source: `Include/arm_nn_types_flt.h:296`
## cmsis_nn_gru_params_f32
`struct` · `c`
```c
struct cmsis_nn_gru_params_f32
```
Parameters for a float32 unidirectional GRU invocation.
GRU has three gates (update, reset, candidate) and, unlike LSTM, no cell state. The hidden state is the layer output.
Source: `Include/arm_nn_types_flt.h:305`
### cmsis_nn_gru_params_f32::time_major
`attribute` · `c`
```c
int32_t time_major
```
Non-zero when input/output tensors are time-major.
Source: `Include/arm_nn_types_flt.h:307`
### cmsis_nn_gru_params_f32::batch_size
`attribute` · `c`
```c
int32_t batch_size
```
Batch size processed per invocation.
Source: `Include/arm_nn_types_flt.h:308`
### cmsis_nn_gru_params_f32::time_steps
`attribute` · `c`
```c
int32_t time_steps
```
Number of time steps processed per invocation.
Source: `Include/arm_nn_types_flt.h:309`
### cmsis_nn_gru_params_f32::input_size
`attribute` · `c`
```c
int32_t input_size
```
Input feature size per time step.
Source: `Include/arm_nn_types_flt.h:310`
### cmsis_nn_gru_params_f32::hidden_size
`attribute` · `c`
```c
int32_t hidden_size
```
Hidden-state size.
Source: `Include/arm_nn_types_flt.h:311`
### cmsis_nn_gru_params_f32::reset_after
`attribute` · `c`
```c
int32_t reset_after
```
Non-zero: reset gate applied after the recurrent matmul (Keras/TFLite default).
Source: `Include/arm_nn_types_flt.h:312`
### cmsis_nn_gru_params_f32::update_gate
`attribute` · `c`
```c
cmsis_nn_gru_gate_f32 update_gate
```
Update gate (z), sigmoid activation.
Source: `Include/arm_nn_types_flt.h:314`
### cmsis_nn_gru_params_f32::reset_gate
`attribute` · `c`
```c
cmsis_nn_gru_gate_f32 reset_gate
```
Reset gate (r), sigmoid activation.
Source: `Include/arm_nn_types_flt.h:315`
### cmsis_nn_gru_params_f32::candidate_gate
`attribute` · `c`
```c
cmsis_nn_gru_gate_f32 candidate_gate
```
Candidate/new gate (n), tanh activation.
Source: `Include/arm_nn_types_flt.h:316`
## cmsis_nn_gru_context_f32
`struct` · `c`
```c
struct cmsis_nn_gru_context_f32
```
Scratch buffers for a float32 GRU invocation.
:::note
`temp1` is sized by `arm_gru_unidirectional_f32_temp1_get_buffer_size()`: `hidden_size` elements when `reset_after == 0` (it holds the reset gate multiplied elementwise by the previous hidden state, r . h_prev). It is unused for the reset-after formulation and may be NULL there. There is no size field and no runtime enforcement: an undersized temp1 is written past on every build target.
:::
:::note
`hidden_state` enables streaming state carry (`batch_size == 1`): when non-NULL it is read as the initial hidden state (seed to zero for a fresh sequence) and overwritten with the final hidden state on return. When NULL the state is zero-initialised and not written back.
:::
Source: `Include/arm_nn_types_flt.h:333`
### cmsis_nn_gru_context_f32::temp1
`attribute` · `c`
```c
float32_t * temp1
```
Scratch required when reset_after == 0; sized by `arm_gru_unidirectional_f32_temp1_get_buffer_size()`.
Source: `Include/arm_nn_types_flt.h:335`
### cmsis_nn_gru_context_f32::hidden_state
`attribute` · `c`
```c
float32_t * hidden_state
```
Optional in/out persistent hidden state [hidden_size] for streaming (batch_size == 1).
Source: `Include/arm_nn_types_flt.h:338`
## cmsis_nn_activation_f16
`struct` · `c`
```c
struct cmsis_nn_activation_f16
```
Activation clamp range for floating-point operators.
Source: `Include/arm_nn_types_flt.h:350`
### cmsis_nn_activation_f16::min
`attribute` · `c`
```c
float16_t min
```
Minimum value used to clamp the result.
Source: `Include/arm_nn_types_flt.h:352`
### cmsis_nn_activation_f16::max
`attribute` · `c`
```c
float16_t max
```
Maximum value used to clamp the result.
Source: `Include/arm_nn_types_flt.h:353`
## cmsis_nn_conv_params_f16
`struct` · `c`
```c
struct cmsis_nn_conv_params_f16
```
Convolution parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:359`
### cmsis_nn_conv_params_f16::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Spatial stride.
Source: `Include/arm_nn_types_flt.h:361`
### cmsis_nn_conv_params_f16::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Spatial zero-padding.
Source: `Include/arm_nn_types_flt.h:362`
### cmsis_nn_conv_params_f16::dilation
`attribute` · `c`
```c
cmsis_nn_tile dilation
```
Spatial dilation.
Source: `Include/arm_nn_types_flt.h:363`
### cmsis_nn_conv_params_f16::activation
`attribute` · `c`
```c
cmsis_nn_activation_f16 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:364`
### cmsis_nn_conv_params_f16::weight_format
`attribute` · `c`
```c
arm_nn_weight_format_flt weight_format
```
Filter storage format.
Source: `Include/arm_nn_types_flt.h:365`
## cmsis_nn_transpose_conv_params_f16
`struct` · `c`
```c
struct cmsis_nn_transpose_conv_params_f16
```
Transpose convolution parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:371`
### cmsis_nn_transpose_conv_params_f16::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Spatial stride.
Source: `Include/arm_nn_types_flt.h:373`
### cmsis_nn_transpose_conv_params_f16::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Spatial zero-padding.
Source: `Include/arm_nn_types_flt.h:374`
### cmsis_nn_transpose_conv_params_f16::padding_offsets
`attribute` · `c`
```c
cmsis_nn_tile padding_offsets
```
Output padding adjustment for transpose convolution.
Source: `Include/arm_nn_types_flt.h:375`
### cmsis_nn_transpose_conv_params_f16::dilation
`attribute` · `c`
```c
cmsis_nn_tile dilation
```
Spatial dilation.
Source: `Include/arm_nn_types_flt.h:376`
### cmsis_nn_transpose_conv_params_f16::activation
`attribute` · `c`
```c
cmsis_nn_activation_f16 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:377`
## cmsis_nn_dw_conv_params_f16
`struct` · `c`
```c
struct cmsis_nn_dw_conv_params_f16
```
Depthwise convolution parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:383`
### cmsis_nn_dw_conv_params_f16::ch_mult
`attribute` · `c`
```c
int32_t ch_mult
```
Channel multiplier. `ch_mult * in_ch = out_ch`.
Source: `Include/arm_nn_types_flt.h:385`
### cmsis_nn_dw_conv_params_f16::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Spatial stride.
Source: `Include/arm_nn_types_flt.h:386`
### cmsis_nn_dw_conv_params_f16::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Spatial zero-padding.
Source: `Include/arm_nn_types_flt.h:387`
### cmsis_nn_dw_conv_params_f16::dilation
`attribute` · `c`
```c
cmsis_nn_tile dilation
```
Spatial dilation.
Source: `Include/arm_nn_types_flt.h:388`
### cmsis_nn_dw_conv_params_f16::activation
`attribute` · `c`
```c
cmsis_nn_activation_f16 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:389`
## cmsis_nn_pool_params_f16
`struct` · `c`
```c
struct cmsis_nn_pool_params_f16
```
Pooling parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:395`
### cmsis_nn_pool_params_f16::stride
`attribute` · `c`
```c
cmsis_nn_tile stride
```
Spatial stride.
Source: `Include/arm_nn_types_flt.h:397`
### cmsis_nn_pool_params_f16::padding
`attribute` · `c`
```c
cmsis_nn_tile padding
```
Spatial zero-padding.
Source: `Include/arm_nn_types_flt.h:398`
### cmsis_nn_pool_params_f16::activation
`attribute` · `c`
```c
cmsis_nn_activation_f16 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:399`
## cmsis_nn_fc_params_f16
`struct` · `c`
```c
struct cmsis_nn_fc_params_f16
```
Fully connected layer parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:405`
### cmsis_nn_fc_params_f16::activation
`attribute` · `c`
```c
cmsis_nn_activation_f16 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:407`
### cmsis_nn_fc_params_f16::weight_format
`attribute` · `c`
```c
arm_nn_weight_format_flt weight_format
```
Weight storage format.
Source: `Include/arm_nn_types_flt.h:408`
## cmsis_nn_bmm_params_f16
`struct` · `c`
```c
struct cmsis_nn_bmm_params_f16
```
Batched matrix multiplication parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:414`
### cmsis_nn_bmm_params_f16::adj_x
`attribute` · `c`
```c
const bool adj_x
```
True when the left-hand-side operand is stored transposed.
Source: `Include/arm_nn_types_flt.h:416`
### cmsis_nn_bmm_params_f16::adj_y
`attribute` · `c`
```c
const bool adj_y
```
True when the right-hand-side operand is stored transposed.
Source: `Include/arm_nn_types_flt.h:417`
### cmsis_nn_bmm_params_f16::activation
`attribute` · `c`
```c
cmsis_nn_activation_f16 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:418`
### cmsis_nn_bmm_params_f16::rhs_format
`attribute` · `c`
```c
arm_nn_weight_format_flt rhs_format
```
Right-hand-side operand storage format. `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` is currently supported only when `adj_x == false` and `adj_y == false`.
Source: `Include/arm_nn_types_flt.h:419`
## cmsis_nn_ew_params_f16
`struct` · `c`
```c
struct cmsis_nn_ew_params_f16
```
Elementwise operator parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:428`
### cmsis_nn_ew_params_f16::activation
`attribute` · `c`
```c
cmsis_nn_activation_f16 activation
```
Output activation clamp range.
Source: `Include/arm_nn_types_flt.h:430`
## cmsis_nn_transpose_params_f16
`struct` · `c`
```c
struct cmsis_nn_transpose_params_f16
```
Transpose parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:436`
### cmsis_nn_transpose_params_f16::num_dims
`attribute` · `c`
```c
int32_t num_dims
```
Number of active dimensions in the permutation.
Source: `Include/arm_nn_types_flt.h:438`
### cmsis_nn_transpose_params_f16::perm
`attribute` · `c`
```c
int32_t perm[4]
```
Permutation indices.
Source: `Include/arm_nn_types_flt.h:439`
### cmsis_nn_transpose_params_f16::layout
`attribute` · `c`
```c
arm_nn_tensor_layout layout
```
Layout convention used to interpret tensor dimensions.
Source: `Include/arm_nn_types_flt.h:440`
## cmsis_nn_svdf_params_f16
`struct` · `c`
```c
struct cmsis_nn_svdf_params_f16
```
Singular value decomposition filter parameters for float32 operators.
Source: `Include/arm_nn_types_flt.h:446`
### cmsis_nn_svdf_params_f16::rank
`attribute` · `c`
```c
int32_t rank
```
SVDF rank.
Source: `Include/arm_nn_types_flt.h:448`
### cmsis_nn_svdf_params_f16::input_activation
`attribute` · `c`
```c
cmsis_nn_activation_f16 input_activation
```
Clamp range applied after the input projection.
Source: `Include/arm_nn_types_flt.h:449`
### cmsis_nn_svdf_params_f16::output_activation
`attribute` · `c`
```c
cmsis_nn_activation_f16 output_activation
```
Clamp range applied to the final output.
Source: `Include/arm_nn_types_flt.h:450`
## cmsis_nn_lstm_gate_f16
`struct` · `c`
```c
struct cmsis_nn_lstm_gate_f16
```
Read-only weights and bias metadata for one float32 LSTM gate.
Source: `Include/arm_nn_types_flt.h:456`
### cmsis_nn_lstm_gate_f16::input_weights
`attribute` · `c`
```c
const float16_t * input_weights
```
Input-to-gate weight matrix.
Source: `Include/arm_nn_types_flt.h:458`
### cmsis_nn_lstm_gate_f16::hidden_weights
`attribute` · `c`
```c
const float16_t * hidden_weights
```
Hidden-state-to-gate weight matrix.
Source: `Include/arm_nn_types_flt.h:459`
### cmsis_nn_lstm_gate_f16::bias
`attribute` · `c`
```c
const float16_t * bias
```
Optional gate bias vector.
Source: `Include/arm_nn_types_flt.h:460`
### cmsis_nn_lstm_gate_f16::activation_type
`attribute` · `c`
```c
arm_nn_activation_type_flt activation_type
```
Gate activation selector.
Source: `Include/arm_nn_types_flt.h:461`
## cmsis_nn_lstm_params_f16
`struct` · `c`
```c
struct cmsis_nn_lstm_params_f16
```
Parameters for a unidirectional float32 LSTM layer.
Source: `Include/arm_nn_types_flt.h:467`
### cmsis_nn_lstm_params_f16::time_major
`attribute` · `c`
```c
int32_t time_major
```
Non-zero when input/output tensors are time-major.
Source: `Include/arm_nn_types_flt.h:469`
### cmsis_nn_lstm_params_f16::batch_size
`attribute` · `c`
```c
int32_t batch_size
```
Batch size processed per invocation.
Source: `Include/arm_nn_types_flt.h:470`
### cmsis_nn_lstm_params_f16::time_steps
`attribute` · `c`
```c
int32_t time_steps
```
Number of time steps processed per invocation.
Source: `Include/arm_nn_types_flt.h:471`
### cmsis_nn_lstm_params_f16::input_size
`attribute` · `c`
```c
int32_t input_size
```
Input feature size per time step.
Source: `Include/arm_nn_types_flt.h:472`
### cmsis_nn_lstm_params_f16::hidden_size
`attribute` · `c`
```c
int32_t hidden_size
```
Hidden-state size.
Source: `Include/arm_nn_types_flt.h:473`
### cmsis_nn_lstm_params_f16::cell_clip
`attribute` · `c`
```c
float16_t cell_clip
```
Optional cell-state clip value.
Source: `Include/arm_nn_types_flt.h:474`
### cmsis_nn_lstm_params_f16::forget_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate_f16 forget_gate
```
Forget gate weights and activation.
Source: `Include/arm_nn_types_flt.h:476`
### cmsis_nn_lstm_params_f16::input_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate_f16 input_gate
```
Input gate weights and activation.
Source: `Include/arm_nn_types_flt.h:477`
### cmsis_nn_lstm_params_f16::cell_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate_f16 cell_gate
```
Cell-update gate weights and activation.
Source: `Include/arm_nn_types_flt.h:478`
### cmsis_nn_lstm_params_f16::output_gate
`attribute` · `c`
```c
cmsis_nn_lstm_gate_f16 output_gate
```
Output gate weights and activation.
Source: `Include/arm_nn_types_flt.h:479`
## cmsis_nn_lstm_context_f16
`struct` · `c`
```c
struct cmsis_nn_lstm_context_f16
```
Scratch and mutable state buffers for a float32 LSTM invocation.
Source: `Include/arm_nn_types_flt.h:485`
### cmsis_nn_lstm_context_f16::temp1
`attribute` · `c`
```c
float16_t * temp1
```
Unused by the current implementation and may be NULL. Sized by `arm_lstm_unidirectional_f16_temp1_get_buffer_size()`, which reports 0; size from the query rather than hard-coding NULL if the buffer is arena-allocated.
Source: `Include/arm_nn_types_flt.h:487`
### cmsis_nn_lstm_context_f16::temp2
`attribute` · `c`
```c
float16_t * temp2
```
Unused by the current implementation and may be NULL. Sized by `arm_lstm_unidirectional_f16_temp2_get_buffer_size()`.
Source: `Include/arm_nn_types_flt.h:490`
### cmsis_nn_lstm_context_f16::cell_state
`attribute` · `c`
```c
float16_t * cell_state
```
Mutable cell-state buffer (in/out when streaming).
Source: `Include/arm_nn_types_flt.h:492`
### cmsis_nn_lstm_context_f16::hidden_state
`attribute` · `c`
```c
float16_t * hidden_state
```
Optional in/out hidden state for streaming; NULL selects stateless. Streaming is NULL-gated, so zero-initialise the context (e.g. designated initialisers) to keep legacy 3-field callers stateless. Matches the quantized `cmsis_nn_lstm_context` contract.
Source: `Include/arm_nn_types_flt.h:493`
## cmsis_nn_gru_gate_f16
`struct` · `c`
```c
struct cmsis_nn_gru_gate_f16
```
Weights and biases for a single float16 GRU gate.
The activation (sigmoid for update/reset, tanh for candidate) is implied by the gate's role and is not stored here. The reset-after formulation keeps the input-projection bias and the recurrent-projection bias separate, because the reset gate multiplies the recurrent projection (including its bias) after the matmul.
Source: `Include/arm_nn_types_flt.h:508`
### cmsis_nn_gru_gate_f16::input_weights
`attribute` · `c`
```c
const float16_t * input_weights
```
Input-to-gate weight matrix [hidden_size, input_size].
Source: `Include/arm_nn_types_flt.h:510`
### cmsis_nn_gru_gate_f16::hidden_weights
`attribute` · `c`
```c
const float16_t * hidden_weights
```
Hidden-to-gate weight matrix [hidden_size, hidden_size].
Source: `Include/arm_nn_types_flt.h:511`
### cmsis_nn_gru_gate_f16::input_bias
`attribute` · `c`
```c
const float16_t * input_bias
```
Optional input-projection bias [hidden_size]. May be NULL.
Source: `Include/arm_nn_types_flt.h:512`
### cmsis_nn_gru_gate_f16::hidden_bias
`attribute` · `c`
```c
const float16_t * hidden_bias
```
Optional recurrent-projection bias [hidden_size]. May be NULL.
Source: `Include/arm_nn_types_flt.h:513`
## cmsis_nn_gru_params_f16
`struct` · `c`
```c
struct cmsis_nn_gru_params_f16
```
Parameters for a float16 unidirectional GRU invocation.
GRU has three gates (update, reset, candidate) and, unlike LSTM, no cell state. The hidden state is the layer output.
Source: `Include/arm_nn_types_flt.h:522`
### cmsis_nn_gru_params_f16::time_major
`attribute` · `c`
```c
int32_t time_major
```
Non-zero when input/output tensors are time-major.
Source: `Include/arm_nn_types_flt.h:524`
### cmsis_nn_gru_params_f16::batch_size
`attribute` · `c`
```c
int32_t batch_size
```
Batch size processed per invocation.
Source: `Include/arm_nn_types_flt.h:525`
### cmsis_nn_gru_params_f16::time_steps
`attribute` · `c`
```c
int32_t time_steps
```
Number of time steps processed per invocation.
Source: `Include/arm_nn_types_flt.h:526`
### cmsis_nn_gru_params_f16::input_size
`attribute` · `c`
```c
int32_t input_size
```
Input feature size per time step.
Source: `Include/arm_nn_types_flt.h:527`
### cmsis_nn_gru_params_f16::hidden_size
`attribute` · `c`
```c
int32_t hidden_size
```
Hidden-state size.
Source: `Include/arm_nn_types_flt.h:528`
### cmsis_nn_gru_params_f16::reset_after
`attribute` · `c`
```c
int32_t reset_after
```
Non-zero: reset gate applied after the recurrent matmul (Keras/TFLite default).
Source: `Include/arm_nn_types_flt.h:529`
### cmsis_nn_gru_params_f16::update_gate
`attribute` · `c`
```c
cmsis_nn_gru_gate_f16 update_gate
```
Update gate (z), sigmoid activation.
Source: `Include/arm_nn_types_flt.h:531`
### cmsis_nn_gru_params_f16::reset_gate
`attribute` · `c`
```c
cmsis_nn_gru_gate_f16 reset_gate
```
Reset gate (r), sigmoid activation.
Source: `Include/arm_nn_types_flt.h:532`
### cmsis_nn_gru_params_f16::candidate_gate
`attribute` · `c`
```c
cmsis_nn_gru_gate_f16 candidate_gate
```
Candidate/new gate (n), tanh activation.
Source: `Include/arm_nn_types_flt.h:533`
## cmsis_nn_gru_context_f16
`struct` · `c`
```c
struct cmsis_nn_gru_context_f16
```
Scratch buffers for a float16 GRU invocation.
:::note
`temp1` is sized by `arm_gru_unidirectional_f16_temp1_get_buffer_size()`: `hidden_size` elements when `reset_after == 0` (it holds the reset gate multiplied elementwise by the previous hidden state, r . h_prev). It is unused for the reset-after formulation and may be NULL there. There is no size field and no runtime enforcement: an undersized temp1 is written past on every build target.
:::
:::note
`hidden_state` enables streaming state carry (`batch_size == 1`): when non-NULL it is read as the initial hidden state (seed to zero for a fresh sequence) and overwritten with the final hidden state on return. When NULL the state is zero-initialised and not written back.
:::
Source: `Include/arm_nn_types_flt.h:550`
### cmsis_nn_gru_context_f16::temp1
`attribute` · `c`
```c
float16_t * temp1
```
Scratch required when reset_after == 0; sized by `arm_gru_unidirectional_f16_temp1_get_buffer_size()`.
Source: `Include/arm_nn_types_flt.h:552`
### cmsis_nn_gru_context_f16::hidden_state
`attribute` · `c`
```c
float16_t * hidden_state
```
Optional in/out persistent hidden state [hidden_size] for streaming (batch_size == 1).
Source: `Include/arm_nn_types_flt.h:555`
## arm_nn_tensor_layout
`enum` · `c`
```c
enum arm_nn_tensor_layout
```
Tensor layout selector for floating-point APIs.
Float public APIs currently accept NHWC layout only.
Source: `Include/arm_nn_types_flt.h:67`
### ARM_NN_LAYOUT_NHWC
`constant` · `c`
```c
ARM_NN_LAYOUT_NHWC = 0
```
Tensor dimensions are ordered as [N, H, W, C].
Source: `Include/arm_nn_types_flt.h:67`
## arm_nn_activation_type
`enum` · `c`
```c
enum arm_nn_activation_type
```
Enum for specifying activation function types
Source: `Include/arm_nn_types.h:78`
### ARM_SIGMOID
`constant` · `c`
```c
ARM_SIGMOID = 0
```
Sigmoid activation function
Source: `Include/arm_nn_types.h:78`
### ARM_TANH
`constant` · `c`
```c
ARM_TANH = 1
```
Tanh activation function
Source: `Include/arm_nn_types.h:78`
## arm_nn_activation_type_flt
`enum` · `c`
```c
enum arm_nn_activation_type_flt
```
Activation selector for floating-point operator APIs.
Numeric values intentionally live in a dedicated floating-point range to avoid overlap with the legacy integer public activation enum.
Source: `Include/arm_nn_types_flt.h:78`
### ARM_NN_FLT_ACT_NONE
`constant` · `c`
```c
ARM_NN_FLT_ACT_NONE = 32
```
Identity activation function.
Source: `Include/arm_nn_types_flt.h:78`
### ARM_NN_FLT_ACT_SIGMOID
`constant` · `c`
```c
ARM_NN_FLT_ACT_SIGMOID = 33
```
Sigmoid activation function.
Source: `Include/arm_nn_types_flt.h:78`
### ARM_NN_FLT_ACT_TANH
`constant` · `c`
```c
ARM_NN_FLT_ACT_TANH = 34
```
Hyperbolic tangent activation function.
Source: `Include/arm_nn_types_flt.h:78`
### ARM_NN_FLT_ACT_RELU
`constant` · `c`
```c
ARM_NN_FLT_ACT_RELU = 35
```
ReLU activation function.
Source: `Include/arm_nn_types_flt.h:78`
### ARM_NN_FLT_ACT_RELU6
`constant` · `c`
```c
ARM_NN_FLT_ACT_RELU6 = 36
```
ReLU6 activation function.
Source: `Include/arm_nn_types_flt.h:78`
### ARM_NN_FLT_ACT_HARDSWISH
`constant` · `c`
```c
ARM_NN_FLT_ACT_HARDSWISH = 37
```
Hard-swish activation function.
Source: `Include/arm_nn_types_flt.h:78`
### ARM_NN_FLT_ACT_LEAKY_RELU
`constant` · `c`
```c
ARM_NN_FLT_ACT_LEAKY_RELU = 38
```
Leaky ReLU activation function.
Source: `Include/arm_nn_types_flt.h:78`
## arm_nn_compare_operation
`enum` · `c`
```c
enum arm_nn_compare_operation
```
Enum for specifying comparison operator
Source: `Include/arm_nn_types.h:85`
### ARM_COMPARE_EQUAL
`constant` · `c`
```c
ARM_COMPARE_EQUAL = 0
```
Returns 1 if lhs == rhs else 0
Source: `Include/arm_nn_types.h:85`
### ARM_COMPARE_NOT_EQUAL
`constant` · `c`
```c
ARM_COMPARE_NOT_EQUAL = 1
```
Returns 1 if lhs != rhs else 0
Source: `Include/arm_nn_types.h:85`
### ARM_COMPARE_GREATER
`constant` · `c`
```c
ARM_COMPARE_GREATER = 2
```
Returns 1 if lhs > rhs else 0
Source: `Include/arm_nn_types.h:85`
### ARM_COMPARE_GREATER_EQUAL
`constant` · `c`
```c
ARM_COMPARE_GREATER_EQUAL = 3
```
Returns 1 if lhs >= rhs else 0
Source: `Include/arm_nn_types.h:85`
### ARM_COMPARE_LESS
`constant` · `c`
```c
ARM_COMPARE_LESS = 4
```
Returns 1 if lhs < rhs else 0
Source: `Include/arm_nn_types.h:85`
### ARM_COMPARE_LESS_EQUAL
`constant` · `c`
```c
ARM_COMPARE_LESS_EQUAL = 5
```
Returns 1 if lhs <= rhs else 0
Source: `Include/arm_nn_types.h:85`
## arm_cmsis_nn_status
`enum` · `c`
```c
enum arm_cmsis_nn_status
```
Function return codes
Source: `Include/arm_nn_types.h:96`
### ARM_CMSIS_NN_SUCCESS
`constant` · `c`
```c
ARM_CMSIS_NN_SUCCESS = 0
```
No error
Source: `Include/arm_nn_types.h:96`
### ARM_CMSIS_NN_ARG_ERROR
`constant` · `c`
```c
ARM_CMSIS_NN_ARG_ERROR = -1
```
One or more arguments are incorrect
Source: `Include/arm_nn_types.h:96`
### ARM_CMSIS_NN_NO_IMPL_ERROR
`constant` · `c`
```c
ARM_CMSIS_NN_NO_IMPL_ERROR = -2
```
No implementation available
Source: `Include/arm_nn_types.h:96`
### ARM_CMSIS_NN_FAILURE
`constant` · `c`
```c
ARM_CMSIS_NN_FAILURE = -3
```
Logical error
Source: `Include/arm_nn_types.h:96`
## arm_nn_dw_kernel_layout_f32
`enum` · `c`
```c
enum arm_nn_dw_kernel_layout_f32
```
Depthwise kernel storage layout selector for floating-point kernels.
Public float depthwise entry points currently use KC storage (`[k][c]`).
Source: `Include/arm_nn_types_flt.h:98`
### ARM_NN_DW_KERNEL_KC
`constant` · `c`
```c
ARM_NN_DW_KERNEL_KC = 0
```
Depthwise kernel stored as `[kernel][channel]`.
Source: `Include/arm_nn_types_flt.h:98`
### ARM_NN_DW_KERNEL_CK
`constant` · `c`
```c
ARM_NN_DW_KERNEL_CK = 1
```
Depthwise kernel stored as `[channel][kernel]`.
Source: `Include/arm_nn_types_flt.h:98`
## arm_nn_weight_format_flt
`enum` · `c`
```c
enum arm_nn_weight_format_flt
```
Weight storage format selector for floating-point operators.
This enum allows frameworks to describe whether weights are provided in the standard public operator layout or in a backend-specific packed layout.
`ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` matches the packed RHS layout consumed by `arm_nn_mat_mult_nt_n_packed_f16/f32`.
The packed NTxN layout exists because MVE kernels typically perform best when output-channel blocks can be loaded contiguously. With the standard `NT x T` formulation, vectorizing over output channels tends to require gather-load accesses to the RHS, which is less efficient than a packed non-transposed RHS layout.
For operators that support `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, supplying offline-repacked constant weights in this layout is therefore generally the preferred way to achieve the best MVE performance.
Source: `Include/arm_nn_types_flt.h:122`
### ARM_NN_WEIGHT_FORMAT_STANDARD
`constant` · `c`
```c
ARM_NN_WEIGHT_FORMAT_STANDARD = 0
```
Standard public operator layout.
Source: `Include/arm_nn_types_flt.h:122`
### ARM_NN_WEIGHT_FORMAT_NT_N_PACKED
`constant` · `c`
```c
ARM_NN_WEIGHT_FORMAT_NT_N_PACKED = 1
```
Packed `[K][N-block]` layout for NTxN matmul helpers.
Source: `Include/arm_nn_types_flt.h:122`
## arm_nn_dw_kernel_layout_f16
`type` · `c`
```c
typedef arm_nn_dw_kernel_layout_f32 arm_nn_dw_kernel_layout_f16
```
Source: `Include/arm_nn_types_flt.h:345`
---
# heliaCORE.groupElementwise
Elementwise add and multiplication functions.
## arm_elementwise_add_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_add_f32(
const float32_t *input_1_vect,
const float32_t *input_2_vect,
float32_t *output,
float32_t out_activation_min,
float32_t out_activation_max,
int32_t block_size
)
```
Elementwise add with optional output clamp.
NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float32_t * | in | Pointer to the first input vector. |
| input_2_vect | const float32_t * | in | Pointer to the second input vector. |
| output | float32_t * | out | Pointer to the output vector. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:714`
## arm_elementwise_sub_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_sub_f32(
const float32_t *input_1_vect,
const float32_t *input_2_vect,
float32_t *output,
float32_t out_activation_min,
float32_t out_activation_max,
int32_t block_size
)
```
Elementwise subtract with optional output clamp.
NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float32_t * | in | Pointer to the first input vector (minuend). |
| input_2_vect | const float32_t * | in | Pointer to the second input vector (subtrahend). |
| output | float32_t * | out | Pointer to the output vector. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:746`
## arm_nn_abs_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_abs_f32(const float32_t *input, float32_t *output, int32_t block_size)
```
Elementwise absolute value.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input vector. |
| output | float32_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:762`
## arm_nn_fill_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_fill_f32(float32_t value, float32_t *output, int32_t block_size)
```
Fill a float32 vector with one value.
Bit copy of `value` into every element (vector splat / plain stores), so a NaN fill value lands bit-exact, sign and payload included. Not named arm_fill_f32: CMSIS-DSP exports that symbol.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| value | float32_t | in | Fill value. |
| output | float32_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to write (0 is a no-op). |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`. |
Source: `Include/arm_nnfunctions_flt.h:778`
## arm_elementwise_mul_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_mul_f32(
const float32_t *input_1_vect,
const float32_t *input_2_vect,
float32_t *output,
float32_t out_activation_min,
float32_t out_activation_max,
int32_t block_size
)
```
Elementwise multiply with optional output clamp.
NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float32_t * | in | Pointer to the first input vector. |
| input_2_vect | const float32_t * | in | Pointer to the second input vector. |
| output | float32_t * | out | Pointer to the output vector. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:805`
## arm_minimum_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_minimum_f32(
const cmsis_nn_context *ctx,
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.
The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float32_t * | in | Input 1, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Input 2, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions_flt.h:858`
## arm_maximum_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_maximum_f32(
const cmsis_nn_context *ctx,
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.
The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float32_t * | in | Input 1, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Input 2, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions_flt.h:894`
## arm_elementwise_sub_broadcast_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_sub_broadcast_f32(
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims,
float32_t out_activation_min,
float32_t out_activation_max
)
```
Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.
Broadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and `output_dims` must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.
Numerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as `scalar - element`, not as the negation of `element - scalar`, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float32_t * | in | Minuend, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Subtrahend, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error. |
Source: `Include/arm_nnfunctions_flt.h:933`
## arm_elementwise_add_broadcast_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_add_broadcast_f32(
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims,
float32_t out_activation_min,
float32_t out_activation_max
)
```
Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.
Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float32_t * | in | First input, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Second input, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |
Source: `Include/arm_nnfunctions_flt.h:957`
## arm_elementwise_mul_broadcast_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_mul_broadcast_f32(
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims,
float32_t out_activation_min,
float32_t out_activation_max
)
```
Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.
Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float32_t * | in | First input, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Second input, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |
Source: `Include/arm_nnfunctions_flt.h:981`
## arm_nn_sqrt_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_sqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)
```
Elementwise square root.
The value path is scalar on every toolchain, because Helium has no vector square root. armclang and ATfE do vectorize the surrounding special-value classification; the results are bit-identical to the GCC scalar build, verified by executing both toolchains' objects (#295). Normal positive inputs evaluate `sqrtf(x)`, which IEEE 754 makes correctly rounded, so results are bit-exact to a float64 reference. Subnormal inputs follow FPSCR.FZ: where flush-to-zero is set - the Corstone-300 FVP default, and any host binary linked at -Ofast, where crtfastmath sets DAZ and FTZ - a subnormal input reads as zero and the result is +0. The float16 pair is immune, because it widens to a normal float32 first. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input vector. |
| output | float32_t * | out | Pointer to the output vector; may alias `input`. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1012`
## arm_rsqrt_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_rsqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)
```
Elementwise reciprocal square root, `1 / sqrt(x)`.
Same value path as arm_nn_sqrt_f32, including its FPSCR.FZ behaviour on subnormal inputs, where the result is +Inf. Normal positive inputs evaluate `1.0f / sqrtf(x)` in float32: two IEEE roundings, so the result is within 1 ulp of the correctly rounded value (measured against a float64 reference; `x = 4^k` is exact). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input vector. |
| output | float32_t * | out | Pointer to the output vector; may alias `input`. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1031`
## arm_elementwise_add_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_add_f16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
float16_t out_activation_min,
float16_t out_activation_max,
int32_t block_size
)
```
Elementwise add with optional output clamp.
NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector. |
| input_2_vect | const float16_t * | in | Pointer to the second input vector. |
| output | float16_t * | out | Pointer to the output vector. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2815`
## arm_elementwise_add_fp16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_add_fp16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
const float16_t out_activation_min,
const float16_t out_activation_max,
const int32_t block_size
)
```
Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate `arm_elementwise_add_f16()`. New code should call `arm_elementwise_add_f16()` instead.
This entry does NOT share the contract of `arm_elementwise_add_f16()`:
:::caution
No argument validation is performed. A NULL `input_1_vect`, `input_2_vect` or `output` is dereferenced rather than reported. A `block_size` of 0 writes nothing and still returns `ARM_CMSIS_NN_SUCCESS`, where `arm_elementwise_add_f16()` returns `ARM_CMSIS_NN_ARG_ERROR`.
:::
:::note
The clamp does not propagate NaN. Both the Helium path (`vminnm`/`vmaxnm`) and the scalar path (the non-propagating `MIN`/`MAX` clamp helper) bound against `out_activation_max` first, so a NaN produced by the addition comes back as `out_activation_max`. `arm_elementwise_add_f16()` documents TensorFlow Lite NaN propagation; this entry does not implement it.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector. Must not be NULL. |
| input_2_vect | const float16_t * | in | Pointer to the second input vector. Must not be NULL. |
| output | float16_t * | out | Pointer to the output vector. Must not be NULL. |
| out_activation_min | const float16_t | in | Minimum output clamp value. |
| out_activation_max | const float16_t | in | Maximum output clamp value. |
| block_size | const int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` unconditionally. |
Source: `Include/arm_nnfunctions_flt.h:2846`
## arm_elementwise_sub_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_sub_f16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
float16_t out_activation_min,
float16_t out_activation_max,
int32_t block_size
)
```
Elementwise subtract with optional output clamp.
NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector (minuend). |
| input_2_vect | const float16_t * | in | Pointer to the second input vector (subtrahend). |
| output | float16_t * | out | Pointer to the output vector. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2856`
## arm_elementwise_squared_difference_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_squared_difference_f16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
int32_t block_size
)
```
Elementwise squared difference of two float16 vectors.
Each output element is calculated as `(input_1_vect[i] - input_2_vect[i])^2`.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector. |
| input_2_vect | const float16_t * | in | Pointer to the second input vector. |
| output | float16_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when an input/output pointer is NULL or `block_size` is less than 1. |
Source: `Include/arm_nnfunctions_flt.h:2876`
## arm_nn_abs_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_abs_f16(const float16_t *input, float16_t *output, int32_t block_size)
```
Elementwise absolute value.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input vector. |
| output | float16_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2884`
## arm_nn_fill_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_fill_f16(float16_t value, float16_t *output, int32_t block_size)
```
Fill a float16 vector with one value; bit copy of `value`, NaN payload included.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| value | float16_t | in | Fill value. |
| output | float16_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to write (0 is a no-op). |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`. |
Source: `Include/arm_nnfunctions_flt.h:2896`
## arm_split_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_split_f16(
const float16_t *input_data,
const int32_t input_dims,
const int32_t *input_shape,
const int32_t axis,
const int32_t num_splits,
const int32_t *split_dims,
float16_t *const *output_data
)
```
Split a float32 tensor of any rank into several tensors along one axis.
Inverse of arm_concatenation_f32; per-split lengths also cover SPLIT_V. Output `s` has the input shape with `input_shape`[axis] replaced by `split_dims`[s]. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input. A dimension of 0 is accepted and copies nothing.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to the flattened (row-major) input. |
| input_dims | const int32_t | in | Number of dimensions in `input_shape` (>= 1). |
| input_shape | const int32_t * | in | Input shape; `input_shape`[axis] must equal the sum of `split_dims`. |
| axis | const int32_t | in | Axis to split along (0 <= axis < input_dims). |
| num_splits | const int32_t | in | Number of outputs (>= 1). |
| split_dims | const int32_t * | in | Array of length `num_splits:` each output's extent along `axis`. |
| output_data | float16_t *const * | out | Array of `num_splits` pointers to the flattened outputs. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, split entry, split sum, NULL pointer or an element count above INT32_MAX. |
Source: `Include/arm_nnfunctions_flt.h:2923`
## arm_strided_slice_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_strided_slice_f16(
const float16_t *input_data,
float16_t *output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const begin_dims,
const cmsis_nn_dims *const stride_dims,
const cmsis_nn_dims *const output_dims
)
```
Strided slice for float32 data (pure copy, TensorFlow Lite compatible).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to input tensor. |
| output_data | float16_t * | out | Pointer to output tensor. |
| input_dims | const cmsis_nn_dims *const | in | Input tensor dimensions. |
| begin_dims | const cmsis_nn_dims *const | in | Begin dimensions for slicing. |
| stride_dims | const cmsis_nn_dims *const | in | Stride dimensions for slicing. |
| output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success. |
Source: `Include/arm_nnfunctions_flt.h:2934`
## arm_elementwise_mul_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_mul_f16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
float16_t out_activation_min,
float16_t out_activation_max,
int32_t block_size
)
```
Elementwise multiply with optional output clamp.
NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector. |
| input_2_vect | const float16_t * | in | Pointer to the second input vector. |
| output | float16_t * | out | Pointer to the output vector. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2944`
## arm_minimum_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_minimum_f16(
const cmsis_nn_context *ctx,
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.
The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float16_t * | in | Input 1, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Input 2, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions_flt.h:2954`
## arm_maximum_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_maximum_f16(
const cmsis_nn_context *ctx,
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.
The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float16_t * | in | Input 1, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Input 2, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions_flt.h:2965`
## arm_elementwise_sub_broadcast_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_sub_broadcast_f16(
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims,
float16_t out_activation_min,
float16_t out_activation_max
)
```
Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.
Broadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and `output_dims` must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.
Numerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as `scalar - element`, not as the negation of `element - scalar`, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.
Half-precision twin: the numerics are those of arm_elementwise_sub_f16 on the materialised operands.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float16_t * | in | Minuend, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Subtrahend, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error. |
Source: `Include/arm_nnfunctions_flt.h:2978`
## arm_elementwise_add_broadcast_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_add_broadcast_f16(
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims,
float16_t out_activation_min,
float16_t out_activation_max
)
```
Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.
Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.
Half-precision twin: the numerics are those of arm_elementwise_add_f16 on the materialised operands.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float16_t * | in | First input, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Second input, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |
Source: `Include/arm_nnfunctions_flt.h:2992`
## arm_elementwise_mul_broadcast_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_mul_broadcast_f16(
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims,
float16_t out_activation_min,
float16_t out_activation_max
)
```
Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.
Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.
Half-precision twin: the numerics are those of arm_elementwise_mul_f16 on the materialised operands.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float16_t * | in | First input, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Second input, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |
Source: `Include/arm_nnfunctions_flt.h:3006`
## arm_nn_sqrt_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_sqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)
```
Elementwise square root of a float16 tensor.
The value path is scalar on every toolchain, because Helium has no vector square root; armclang and ATfE vectorize the surrounding classification into an MVE loop and produce bit-identical results, verified by executing their objects (#295). Each element is widened to float32, `sqrtf` is evaluated there and the result is rounded once to float16. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 square root (0 ulp, #295). Widening first also makes this pair immune to FPSCR.FZ, which flushes float32 subnormals in the f32 pair. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. |
| output | float16_t * | out | Pointer to the output tensor; may alias `input`. |
| block_size | int32_t | in | Number of tensor elements. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3036`
## arm_rsqrt_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_rsqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)
```
Elementwise reciprocal square root of a float16 tensor, `1 / sqrt(x)`.
Same value path as arm_nn_sqrt_f16: widen to float32, evaluate `1.0f / sqrtf(x)` there, round once to float16, and so also immune to FPSCR.FZ. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 reciprocal square root (0 ulp, #295). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. |
| output | float16_t * | out | Pointer to the output tensor; may alias `input`. |
| block_size | int32_t | in | Number of tensor elements. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3055`
---
# heliaCORE.groupSupport
Internal Support functions. Not intended to be called direclty by a CMSIS-NN user.
## arm_nn_exp_poly_coeffs_f32
`attribute` · `c`
```c
const float32_t arm_nn_exp_poly_coeffs_f32[8]
```
Polynomial coefficients used by the float32 MVE exp approximation.
Source: `Include/arm_nnsupportfunctions_flt.h:70`
## arm_nn_exp2_lut_f32
`attribute` · `c`
```c
const float32_t arm_nn_exp2_lut_f32[257]
```
LUT for `2^(i/256)` used by the float32 LUT softmax approximation.
Stores 257 samples for `i = 0..256` so interpolation can safely read `lut[idx + 1]` while indexing the 256 fractional segments.
Source: `Include/arm_nnsupportfunctions_flt.h:78`
## arm_nn_softmax_floor_to_int_f32
`function` · `c`
```c
static int32_t arm_nn_softmax_floor_to_int_f32(float32_t x)
```
Floor of `x` as an int32_t.
Precondition: `x` must already be reduced to the int32_t range and must not be NaN the float-to-int conversion below is undefined otherwise. The only caller, arm_nn_softmax_exp_lut_f32(), guarantees this by clamping its input to [-80, 80] (NaN included, see there) before scaling by log2(e), which bounds `x` to +/-116.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | in | Value to floor. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Largest int32_t not greater than `x`. |
Source: `Include/arm_nnsupportfunctions_flt.h:92`
## arm_nn_softmax_fp32_from_bits
`function` · `c`
```c
static float32_t arm_nn_softmax_fp32_from_bits(uint32_t bits)
```
Reinterpret a 32-bit pattern as a float32.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bits | uint32_t | in | IEEE-754 binary32 bit pattern. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The float32 value with the bit pattern `bits`. |
Source: `Include/arm_nnsupportfunctions_flt.h:104`
## arm_nn_softmax_exp2i_f32
`function` · `c`
```c
static float32_t arm_nn_softmax_exp2i_f32(int32_t n)
```
Compute `2^n` as a float32 by building the exponent field directly.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| n | int32_t | in | Integer exponent. Clamped to the normal float32 exponent range `[-126, 127]`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `2^n` as a float32. |
Source: `Include/arm_nnsupportfunctions_flt.h:121`
## arm_nn_softmax_exp_taylor_f32
`function` · `c`
```c
static float32_t arm_nn_softmax_exp_taylor_f32(float32_t x)
```
Taylor/Estrin exp approximation for float32 softmax helpers.
The polynomial is evaluated on r in [-ln2/2, ln2/2]. Coefficients come from the Maclaurin series of exp(r): exp(r) ~= 1 + r + r^2/2! + r^3/3! + r^4/4! + r^5/5! + r^6/6! Grouped via Estrin to reduce dependency depth: p = (1 + r) + r^2*(1/2 + r/6) + r^4*(1/24 + r/120) + r^6*(1/720)
Range reduction follows: x = n * ln(2) + r, exp(x) = exp(r) * 2^n
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | in | Exponent argument. Clamped to `[-80, 80]` before evaluation. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Approximation of `exp(x)`. |
Source: `Include/arm_nnsupportfunctions_flt.h:147`
## arm_nn_softmax_exp_lut_f32
`function` · `c`
```c
static float32_t arm_nn_softmax_exp_lut_f32(float32_t x)
```
LUT-based exp approximation for float32 softmax helpers.
Splits `x * log2(e)` into an integer part handled by arm_nn_softmax_exp2i_f32() and a fractional part interpolated linearly from `arm_nn_exp2_lut_f32`.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | in | Exponent argument. Clamped to `[-80, 80]` before evaluation; NaN is flushed to `80`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Approximation of `exp(x)`. |
Source: `Include/arm_nnsupportfunctions_flt.h:184`
## arm_nn_softmax_exp_scalar_f32
`function` · `c`
```c
static float32_t arm_nn_softmax_exp_scalar_f32(float32_t x)
```
Scalar exp approximation used by the float32 softmax paths.
Dispatches to arm_nn_softmax_exp_taylor_f32() when `ARM_NN_USE_EXP_TAYLOR` is defined and to arm_nn_softmax_exp_lut_f32() otherwise.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | in | Exponent argument. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Approximation of `exp(x)`. |
Source: `Include/arm_nnsupportfunctions_flt.h:254`
## arm_nn_tanh_lut_f32
`attribute` · `c`
```c
const float32_t arm_nn_tanh_lut_f32[385]
```
LUT for tanh(x) sampled over `x in [0, 6]` for float32 helpers.
Stores 385 samples so interpolation can safely read `lut[idx + 1]` while indexing the 384 fractional segments across the interval. The grid spacing (`6/384 == 1/64`) matches the earlier 257-entry `[0, 4]` table, so entries `0..256` are bit-identical to it and the index multiplier is unchanged. Generated by `scripts/gen_tanh_lut_f32.py`.
Source: `Include/arm_nnsupportfunctions_flt.h:276`
## arm_memcpy_f32
`function` · `c`
```c
static void arm_memcpy_f32(float32_t *dst, const float32_t *src, uint32_t block_size)
```
Copy a float32 vector.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | float32_t * | out | Destination buffer. |
| src | const float32_t * | in | Source buffer. |
| block_size | uint32_t | in | Number of elements to copy. |
Source: `Include/arm_nnsupportfunctions_flt.h:311`
## arm_memset_f32
`function` · `c`
```c
static void arm_memset_f32(float32_t *dst, const float32_t val, uint32_t block_size)
```
Set a float32 vector to a constant value.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | float32_t * | out | Destination buffer. |
| val | const float32_t | in | Fill value. |
| block_size | uint32_t | in | Number of elements to write. |
Source: `Include/arm_nnsupportfunctions_flt.h:334`
## arm_nn_depthwise_conv1d_k3_nhwc_f32
`function` · `c`
```c
void arm_nn_depthwise_conv1d_k3_nhwc_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel,
const float32_t *b,
float32_t *out,
int32_t out_w
)
```
Specialized NHWC depthwise 1D kernel for `k=3`, `ch_mult=1` (float32).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input (and output) channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float32_t * | in | Depthwise weights with shape `[3][in_c]`. |
| b | const float32_t * | in | Optional bias vector of `in_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:367`
## arm_nn_conv1d_k5_nhwc_f32
`function` · `c`
```c
void arm_nn_conv1d_k5_nhwc_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel,
const float32_t *b,
float32_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=5` (float32).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float32_t * | in | Weights with shape `[out_c][5][in_c]`. |
| b | const float32_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |
Source: `Include/arm_nnsupportfunctions_flt.h:387`
## arm_nn_conv1d_k5_packed_f32
`function` · `c`
```c
void arm_nn_conv1d_k5_packed_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel_packed,
const float32_t *b,
float32_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=5` (float32, packed weights).
The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f32`, i.e. `[(5 * in_c)][out_c_block_of_4]`.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float32_t * | in | Weights packed in output-channel blocks of 4 as described above. |
| b | const float32_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |
Source: `Include/arm_nnsupportfunctions_flt.h:411`
## arm_nn_conv1d_k3_nhwc_f32
`function` · `c`
```c
void arm_nn_conv1d_k3_nhwc_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel,
const float32_t *b,
float32_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=3` (float32).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float32_t * | in | Weights with shape `[out_c][3][in_c]`. |
| b | const float32_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:432`
## arm_nn_conv1d_k3_packed_f32
`function` · `c`
```c
void arm_nn_conv1d_k3_packed_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel_packed,
const float32_t *b,
float32_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=3` (float32, packed weights).
The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f32`, i.e. `[(3 * in_c)][out_c_block_of_4]`.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float32_t * | in | Weights packed in output-channel blocks of 4 as described above. |
| b | const float32_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:456`
## arm_nn_maxpool1d_k3s3_nhwc_f32
`function` · `c`
```c
void arm_nn_maxpool1d_k3s3_nhwc_f32(const float32_t *x_nhwc, int32_t in_c, int32_t in_w, float32_t *out, int32_t out_w)
```
Specialized NHWC max-pool 1D kernel for `k=3`, `s=3` (float32).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `3*ow..3*ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:474`
## arm_nn_maxpool1d_k2s2_nhwc_noclip_f32
`function` · `c`
```c
void arm_nn_maxpool1d_k2s2_nhwc_noclip_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
float32_t *out,
int32_t out_w
)
```
Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` without output clamp (float32).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `2*ow..2*ow+1`. |
Source: `Include/arm_nnsupportfunctions_flt.h:489`
## arm_nn_maxpool1d_k2s2_nhwc_f32
`function` · `c`
```c
void arm_nn_maxpool1d_k2s2_nhwc_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
float32_t *out,
int32_t out_w,
float32_t act_min,
float32_t act_max
)
```
Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` with clamp (float32).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `2*ow..2*ow+1`. |
| act_min | float32_t | in | Lower clamp bound applied to `out`. |
| act_max | float32_t | in | Upper clamp bound applied to `out`. |
Source: `Include/arm_nnsupportfunctions_flt.h:506`
## arm_nn_mat_mult_nt_t_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f32(
const float32_t *lhs,
const float32_t *rhs,
const float32_t *bias,
float32_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float32_t activation_min,
float32_t activation_max
)
```
Matrix multiply with non-transposed lhs and transposed rhs rows (float32).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float32_t * | in | Left-hand matrix stored row-major. |
| rhs | const float32_t * | in | Right-hand matrix stored row-major, one row per output channel. |
| bias | const float32_t * | in | Optional bias vector. |
| dst | float32_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of rows in `rhs`. |
| rhs_cols | int32_t | in | Number of columns in `rhs`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float32_t | in | Lower clamp bound. |
| activation_max | float32_t | in | Upper clamp bound. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnsupportfunctions_flt.h:529`
## arm_nn_mat_mult_nt_n_packed_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f32(
const float32_t *lhs,
const float32_t *rhs_packed,
const float32_t *bias,
float32_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float32_t activation_min,
float32_t activation_max
)
```
Matrix multiply with non-transposed lhs and packed non-transposed rhs (float32).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float32_t * | in | Left-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`. |
| rhs_packed | const float32_t * | in | Right-hand matrix with logical shape `[rhs_cols, rhs_rows]`, packed in column blocks of 4. The final block uses the same packed stride and inactive tail lanes are ignored. |
| bias | const float32_t * | in | Optional bias vector. |
| dst | float32_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of logical output columns in the unpacked rhs matrix. |
| rhs_cols | int32_t | in | Shared reduction dimension `K`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float32_t | in | Lower clamp bound. |
| activation_max | float32_t | in | Upper clamp bound. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnsupportfunctions_flt.h:556`
## arm_nn_pack_conv_patch_f32
`function` · `c`
```c
void arm_nn_pack_conv_patch_f32(
const float32_t *input,
int32_t in_h,
int32_t in_w,
int32_t in_c,
int32_t kernel_h,
int32_t kernel_w,
int32_t stride_h,
int32_t stride_w,
int32_t pad_h,
int32_t pad_w,
int32_t dilation_h,
int32_t dilation_w,
int32_t out_y,
int32_t out_x,
float32_t pad_value,
float32_t *patch_row
)
```
Pack a single convolution patch into one row of a contiguous float32 patch matrix.
Developers familiar with im2row/im2col terminology can think of this as packing one output patch into one row.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Input tensor for one batch in NHWC layout with shape `[in_h][in_w][in_c]`. |
| in_h | int32_t | in | Input height. |
| in_w | int32_t | in | Input width. |
| in_c | int32_t | in | Number of input channels. |
| kernel_h | int32_t | in | Kernel height. |
| kernel_w | int32_t | in | Kernel width. |
| stride_h | int32_t | in | Vertical stride. |
| stride_w | int32_t | in | Horizontal stride. |
| pad_h | int32_t | in | Top padding. |
| pad_w | int32_t | in | Left padding. |
| dilation_h | int32_t | in | Vertical dilation. |
| dilation_w | int32_t | in | Horizontal dilation. |
| out_y | int32_t | in | Output row index of the patch to pack. |
| out_x | int32_t | in | Output column index of the patch to pack. |
| pad_value | float32_t | in | Value written for taps that fall outside the input. |
| patch_row | float32_t * | out | Destination row of `kernel_h * kernel_w * in_c` elements, ordered `[kernel_h][kernel_w][in_c]`. |
Source: `Include/arm_nnsupportfunctions_flt.h:590`
## arm_nn_softmax_1x2_f32
`function` · `c`
```c
void arm_nn_softmax_1x2_f32(const float32_t *in, float32_t *out)
```
Specialized softmax helper for a single float32 row of length 2.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in | const float32_t * | in | Pointer to two contiguous float32 input values. |
| out | float32_t * | out | Pointer to two contiguous float32 output values. |
Source: `Include/arm_nnsupportfunctions_flt.h:613`
## ARM_NN_F16_ACC_BLOCK
`macro` · `c`
```c
#define ARM_NN_F16_ACC_BLOCK (32)
```
Blockwise float16 accumulation on the MVE legs (AmbiqAI/ns-cmsis-nn#586).
A float16 accumulator lane sums at most ARM_NN_F16_ACC_BLOCK taps, in the kernel's tap order, before its partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once. The `_acc16` entries instantiate the same kernel bodies with ARM_NN_F16_ACC_BLOCK_NONE, which never folds.
Source: `Include/arm_nnsupportfunctions_flt.h:626`
## ARM_NN_F16_ACC_BLOCK_NONE
`macro` · `c`
```c
#define ARM_NN_F16_ACC_BLOCK_NONE (INT32_MAX)
```
Source: `Include/arm_nnsupportfunctions_flt.h:627`
## arm_nn_exp_poly_coeffs_f16
`attribute` · `c`
```c
const float32_t arm_nn_exp_poly_coeffs_f16[8]
```
Polynomial coefficients used by the float16 MVE exp approximation.
The float16 MVE helper evaluates the polynomial in widened float32 lanes, but it uses a dedicated coefficient table to keep the float16 path isolated from the float32 feature gate and softmax support stack.
Source: `Include/arm_nnsupportfunctions_flt.h:636`
## arm_nn_exp2_lut_f16
`attribute` · `c`
```c
const uint16_t arm_nn_exp2_lut_f16[257]
```
Quantized binary16 LUT for `2^(i/256)` used by float16 helpers.
Stores 257 samples for `i = 0..256` so interpolation can safely read `lut[idx + 1]` while indexing the 256 fractional segments.
Source: `Include/arm_nnsupportfunctions_flt.h:644`
## arm_nn_tanh_lut_f16
`attribute` · `c`
```c
const uint16_t arm_nn_tanh_lut_f16[257]
```
Quantized binary16 LUT for tanh(x) with `x in [0, 4]`.
Stores 257 samples so interpolation can safely read `lut[idx + 1]` while indexing the 256 fractional segments across the interval.
Source: `Include/arm_nnsupportfunctions_flt.h:652`
## arm_nn_softmax_fp16_from_bits
`function` · `c`
```c
static float16_t arm_nn_softmax_fp16_from_bits(uint16_t bits)
```
Reinterpret a 16-bit pattern as a float16.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bits | uint16_t | in | IEEE-754 binary16 bit pattern. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The float16 value with the bit pattern `bits`. |
Source: `Include/arm_nnsupportfunctions_flt.h:660`
## arm_nn_softmax_floor_to_int_f16
`function` · `c`
```c
static int32_t arm_nn_softmax_floor_to_int_f16(float16_t x)
```
Floor of `x` as an int32_t.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | in | Value to floor. Must be finite and within the int32_t range. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Largest int32_t not greater than `x`. |
Source: `Include/arm_nnsupportfunctions_flt.h:677`
## arm_nn_softmax_exp2i_f16
`function` · `c`
```c
static float16_t arm_nn_softmax_exp2i_f16(int32_t n)
```
Compute `2^n` as a float16 by building the exponent field directly.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| n | int32_t | in | Integer exponent. Clamped to the normal float16 exponent range `[-14, 15]`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `2^n` as a float16. |
Source: `Include/arm_nnsupportfunctions_flt.h:690`
## arm_nn_softmax_exp_taylor_f16
`function` · `c`
```c
static float16_t arm_nn_softmax_exp_taylor_f16(float16_t x)
```
Taylor/Estrin exp approximation for float16 softmax helpers.
The evaluation uses float32 intermediates to keep the approximation stable, but it is fully independent from the float32 softmax support tables.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | in | Exponent argument. Clamped to `[-80, 80]` before evaluation. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Approximation of `exp(x)`. |
Source: `Include/arm_nnsupportfunctions_flt.h:710`
## arm_nn_softmax_exp_lut_f16
`function` · `c`
```c
static float16_t arm_nn_softmax_exp_lut_f16(float16_t x)
```
LUT-based exp approximation for float16 softmax helpers.
Splits `x * log2(e)` into an integer part handled by arm_nn_softmax_exp2i_f16() and a fractional part interpolated linearly from `arm_nn_exp2_lut_f16`, using float32 intermediates.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | in | Exponent argument. Clamped to `[-80, 80]` before evaluation. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Approximation of `exp(x)`. |
Source: `Include/arm_nnsupportfunctions_flt.h:746`
## arm_nn_softmax_exp_scalar_f16
`function` · `c`
```c
static float16_t arm_nn_softmax_exp_scalar_f16(float16_t x)
```
Scalar exp approximation used by the float16 softmax paths.
Dispatches to arm_nn_softmax_exp_taylor_f16() when `ARM_NN_USE_EXP_TAYLOR` is defined and to arm_nn_softmax_exp_lut_f16() otherwise.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | in | Exponent argument. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Approximation of `exp(x)`. |
Source: `Include/arm_nnsupportfunctions_flt.h:787`
## arm_memcpy_f16
`function` · `c`
```c
static void arm_memcpy_f16(float16_t *dst, const float16_t *src, uint32_t block_size)
```
Copy a float16 vector.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | float16_t * | out | Destination buffer. |
| src | const float16_t * | in | Source buffer. |
| block_size | uint32_t | in | Number of elements to copy. |
Source: `Include/arm_nnsupportfunctions_flt.h:979`
## arm_memset_f16
`function` · `c`
```c
static void arm_memset_f16(float16_t *dst, const float16_t val, uint32_t block_size)
```
Set a float16 vector to a constant value.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | float16_t * | out | Destination buffer. |
| val | const float16_t | in | Fill value. |
| block_size | uint32_t | in | Number of elements to write. |
Source: `Include/arm_nnsupportfunctions_flt.h:1002`
## arm_nn_depthwise_conv2x5_nhwc_f16
`function` · `c`
```c
void arm_nn_depthwise_conv2x5_nhwc_f16(
const float16_t *x_nhwc,
int32_t batches,
int32_t in_c,
int32_t in_w,
int32_t ch_mult,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_w,
float16_t act_min,
float16_t act_max
)
```
Specialized NHWC depthwise `2x5` kernel (float16).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input tensor in NHWC layout with shape `[batches][2][in_w][in_c]`. |
| batches | int32_t | in | Number of batches. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. |
| ch_mult | int32_t | in | Channel multiplier; the output has `in_c * ch_mult` channels. |
| kernel | const float16_t * | in | Depthwise weights with shape `[2][5][in_c * ch_mult]`. |
| b | const float16_t * | in | Optional bias vector of `in_c * ch_mult` elements. May be NULL. |
| out | float16_t * | out | Output tensor in NHWC layout with shape `[batches][1][out_w][in_c * ch_mult]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input columns `ow..ow+4`. |
| act_min | float16_t | in | Lower clamp bound applied to `out`. |
| act_max | float16_t | in | Upper clamp bound applied to `out`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1039`
## arm_nn_depthwise_conv1d_k3_nhwc_f16
`function` · `c`
```c
void arm_nn_depthwise_conv1d_k3_nhwc_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_w
)
```
Specialized NHWC depthwise 1D kernel for `k=3`, `ch_mult=1` (float32).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input (and output) channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Depthwise weights with shape `[3][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `in_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1054`
## arm_nn_conv1d_k5_nhwc_f16
`function` · `c`
```c
void arm_nn_conv1d_k5_nhwc_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=5` (float32).
:::note
MVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 5 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 6 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Weights with shape `[out_c][5][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1073`
## arm_nn_conv1d_k5_nhwc_f16_acc16
`function` · `c`
```c
void arm_nn_conv1d_k5_nhwc_f16_acc16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=5` (float32).
:::note
MVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 5 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 6 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).
:::
:::note
Every MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k5_nhwc_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Weights with shape `[out_c][5][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1088`
## arm_nn_conv1d_k5_packed_f16
`function` · `c`
```c
void arm_nn_conv1d_k5_packed_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel_packed,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=5` (float16, packed weights).
The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(5 * in_c)][out_c_block_of_8]`.
:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 6 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 6) this is the float16-lane result.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float16_t * | in | Weights packed in output-channel blocks of 8 as described above. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1118`
## arm_nn_conv1d_k5_packed_f16_acc16
`function` · `c`
```c
void arm_nn_conv1d_k5_packed_f16_acc16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel_packed,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=5` (float16, packed weights).
The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(5 * in_c)][out_c_block_of_8]`.
:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 6 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 6) this is the float16-lane result.
:::
:::note
Every MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k5_packed_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float16_t * | in | Weights packed in output-channel blocks of 8 as described above. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1133`
## arm_nn_conv1d_k3_nhwc_f16
`function` · `c`
```c
void arm_nn_conv1d_k3_nhwc_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=3` (float32).
:::note
MVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 3 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 10 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Weights with shape `[out_c][3][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1153`
## arm_nn_conv1d_k3_nhwc_f16_acc16
`function` · `c`
```c
void arm_nn_conv1d_k3_nhwc_f16_acc16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=3` (float32).
:::note
MVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 3 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 10 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).
:::
:::note
Every MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k3_nhwc_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Weights with shape `[out_c][3][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1168`
## arm_nn_conv1d_k3_packed_f16
`function` · `c`
```c
void arm_nn_conv1d_k3_packed_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel_packed,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=3` (float16, packed weights).
The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(3 * in_c)][out_c_block_of_8]`.
:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 10 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 10) this is the float16-lane result.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float16_t * | in | Weights packed in output-channel blocks of 8 as described above. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1198`
## arm_nn_conv1d_k3_packed_f16_acc16
`function` · `c`
```c
void arm_nn_conv1d_k3_packed_f16_acc16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel_packed,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)
```
Specialized NHWC 1D convolution kernel for `k=3` (float16, packed weights).
The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(3 * in_c)][out_c_block_of_8]`.
:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 10 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 10) this is the float16-lane result.
:::
:::note
Every MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k3_packed_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float16_t * | in | Weights packed in output-channel blocks of 8 as described above. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1213`
## arm_nn_maxpool1d_k3s3_nhwc_f16
`function` · `c`
```c
void arm_nn_maxpool1d_k3s3_nhwc_f16(const float16_t *x_nhwc, int32_t in_c, int32_t in_w, float16_t *out, int32_t out_w)
```
Specialized NHWC max-pool 1D kernel for `k=3`, `s=3` (float16).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `3*ow..3*ow+2`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1227`
## arm_nn_maxpool1d_k2s2_nhwc_noclip_f16
`function` · `c`
```c
void arm_nn_maxpool1d_k2s2_nhwc_noclip_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
float16_t *out,
int32_t out_w
)
```
Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` without output clamp (float16).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `2*ow..2*ow+1`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1238`
## arm_nn_maxpool1d_k2s2_nhwc_f16
`function` · `c`
```c
void arm_nn_maxpool1d_k2s2_nhwc_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
float16_t *out,
int32_t out_w,
float16_t act_min,
float16_t act_max
)
```
Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` with clamp (float16).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `2*ow..2*ow+1`. |
| act_min | float16_t | in | Lower clamp bound applied to `out`. |
| act_max | float16_t | in | Upper clamp bound applied to `out`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1249`
## arm_nn_mat_mult_nt_t_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f16(
const float16_t *lhs,
const float16_t *rhs,
const float16_t *bias,
float16_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float16_t activation_min,
float16_t activation_max
)
```
Matrix multiply with non-transposed lhs and transposed rhs rows (float32).
:::note
Accumulation width per leg. MVE legs accumulate blockwise (AmbiqAI/ns-cmsis-nn#586, superseding the float16-lane choice of #417 / #446 for the MVE legs). Up to rhs_cols 32 nothing changes: per-k float16 lanes on the gather path (rhs_cols below the contiguous-K threshold), lane-partial sums then one float16 reduction plus the bias in float16 at rhs_cols 32. Above 32, on the contiguous-K path and the remainder rows, each lane (element k goes to lane k % 8) sums at most 32 of its own taps (256 elements) in float16; each block's lanes are then widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (set by the first block, added to by later ones); the four pair accumulators are summed once as (0+1) + (2+3), so a single block sums ((0+1) + (2+3)) + ((4+5) + (6+7)), the bias is added in float32 and the total rounds to float16 once before the clamp. The float16 reduction up to rhs_cols 32 is ordered by the compiler, which may reorder it under -ffast-math; only the fold's order is fixed. arm_nn_mat_mult_nt_t_f16_acc16 keeps the float16 lanes and float16 reduction throughout. The scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates bias and every product in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457).
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float16_t * | in | Left-hand matrix stored row-major. |
| rhs | const float16_t * | in | Right-hand matrix stored row-major, one row per output channel. |
| bias | const float16_t * | in | Optional bias vector. |
| dst | float16_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of rows in `rhs`. |
| rhs_cols | int32_t | in | Number of columns in `rhs`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float16_t | in | Lower clamp bound. |
| activation_max | float16_t | in | Upper clamp bound. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnsupportfunctions_flt.h:1273`
## arm_nn_mat_mult_nt_t_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f16_acc16(
const float16_t *lhs,
const float16_t *rhs,
const float16_t *bias,
float16_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float16_t activation_min,
float16_t activation_max
)
```
arm_nn_mat_mult_nt_t_f16 with every MVE accumulator lane in float16 (no blockwise fold).
Same arguments, return codes and scalar leg as arm_nn_mat_mult_nt_t_f16; see its accumulation note.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float16_t * | in | Left-hand matrix, row-major `[lhs_rows, rhs_cols]`. |
| rhs | const float16_t * | in | Right-hand matrix, row-major `[rhs_rows, rhs_cols]` (transposed operand). |
| bias | const float16_t * | in | Optional bias vector of `rhs_rows` elements. |
| dst | float16_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of rows in `rhs`. |
| rhs_cols | int32_t | in | Shared reduction dimension `K`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float16_t | in | Lower clamp bound. |
| activation_max | float16_t | in | Upper clamp bound. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnsupportfunctions_flt.h:1301`
## arm_nn_mat_mult_nt_n_packed_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f16(
const float16_t *lhs,
const float16_t *rhs_packed,
const float16_t *bias,
float16_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float16_t activation_min,
float16_t activation_max
)
```
Matrix multiply with non-transposed lhs and packed non-transposed rhs (float16).
:::note
On non-MVE builds the output clamp is the bit-classified scalar clamp of #380, so a NaN accumulator (a NaN in `lhs`, `rhs_packed` or `bias`) propagates to `dst` at every optimization level on the gated toolchains, including the shipped -Ofast. On MVE builds the clamp is vmaxnmq/vminnmq with no NaN restore, so a NaN resolves to a clamp bound there instead.
:::
:::note
Accumulation width per leg: the MVE leg accumulates blockwise (AmbiqAI/ns-cmsis-nn#586): one lane per output column, per-k, the bias opening the first block; every 32 k the float16 partial is widened exactly into per-lane float32 accumulators, which round to float16 once before the clamp (rhs_cols up to 32: exactly the float16-lane result; arm_nn_mat_mult_nt_n_packed_f16_acc16 keeps float16 lanes throughout). The scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates bias and every product in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457).
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float16_t * | in | Left-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`. |
| rhs_packed | const float16_t * | in | Right-hand matrix with logical shape `[rhs_cols, rhs_rows]`, packed in column blocks of 8. The final block uses the same packed stride and inactive tail lanes are ignored. |
| bias | const float16_t * | in | Optional bias vector. |
| dst | float16_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of logical output columns in the unpacked rhs matrix. |
| rhs_cols | int32_t | in | Shared reduction dimension `K`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float16_t | in | Lower clamp bound. |
| activation_max | float16_t | in | Upper clamp bound. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnsupportfunctions_flt.h:1342`
## arm_nn_mat_mult_nt_n_packed_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f16_acc16(
const float16_t *lhs,
const float16_t *rhs_packed,
const float16_t *bias,
float16_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float16_t activation_min,
float16_t activation_max
)
```
arm_nn_mat_mult_nt_n_packed_f16 with every MVE accumulator lane in float16 (no blockwise fold).
Same arguments, return codes and scalar leg as arm_nn_mat_mult_nt_n_packed_f16; see its accumulation note.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float16_t * | in | Left-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`. |
| rhs_packed | const float16_t * | in | Right-hand matrix packed in column blocks of 8. |
| bias | const float16_t * | in | Optional bias vector. |
| dst | float16_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of logical output columns in the unpacked rhs matrix. |
| rhs_cols | int32_t | in | Shared reduction dimension `K`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float16_t | in | Lower clamp bound. |
| activation_max | float16_t | in | Upper clamp bound. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnsupportfunctions_flt.h:1370`
## arm_nn_lstm_step_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_lstm_step_f16(
const float16_t *data_in,
const float16_t *hidden_in,
float16_t *hidden_out,
const cmsis_nn_lstm_params_f16 *params,
cmsis_nn_lstm_context_f16 *buffers,
const int32_t batch_offset
)
```
Update LSTM function for an iteration step using float16 input, output and state.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const float16_t * | in | Data input pointer. |
| hidden_in | const float16_t * | in | Hidden state / recurrent input pointer. May be NULL for the first step. |
| hidden_out | float16_t * | out | Hidden state / recurrent output pointer. |
| params | const cmsis_nn_lstm_params_f16 * | in | Struct containing all information about the LSTM operator. |
| buffers | cmsis_nn_lstm_context_f16 * | in, out | Struct containing pointers to mutable cell-state storage. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params/buffers or buffers->cell_state, batch_offset <= 0). |
Source: `Include/arm_nnsupportfunctions_flt.h:1394`
## arm_nn_gru_step_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_gru_step_f16(
const float16_t *data_in,
const float16_t *hidden_in,
float16_t *hidden_out,
const cmsis_nn_gru_params_f16 *params,
cmsis_nn_gru_context_f16 *buffers,
const int32_t batch_offset
)
```
Update GRU function for a single iteration step using float16 data.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const float16_t * | in | Data input pointer for this time step. |
| hidden_in | const float16_t * | in | Recurrent input pointer. NULL for the first step (h_prev = 0). |
| hidden_out | float16_t * | out | Hidden-state output pointer for this time step. |
| params | const cmsis_nn_gru_params_f16 * | in | Struct describing the GRU operator. |
| buffers | cmsis_nn_gru_context_f16 * | in, out | Scratch buffers. temp1 (>= hidden_size) is required when reset_after == 0. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params, batch_offset <= 0, or missing temp1 when reset_after == 0). |
Source: `Include/arm_nnsupportfunctions_flt.h:1414`
## arm_nn_pack_conv_patch_f16
`function` · `c`
```c
void arm_nn_pack_conv_patch_f16(
const float16_t *input,
int32_t in_h,
int32_t in_w,
int32_t in_c,
int32_t kernel_h,
int32_t kernel_w,
int32_t stride_h,
int32_t stride_w,
int32_t pad_h,
int32_t pad_w,
int32_t dilation_h,
int32_t dilation_w,
int32_t out_y,
int32_t out_x,
float16_t pad_value,
float16_t *patch_row
)
```
Pack a single convolution patch into one row of a contiguous float32 patch matrix.
Developers familiar with im2row/im2col terminology can think of this as packing one output patch into one row.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Input tensor for one batch in NHWC layout with shape `[in_h][in_w][in_c]`. |
| in_h | int32_t | in | Input height. |
| in_w | int32_t | in | Input width. |
| in_c | int32_t | in | Number of input channels. |
| kernel_h | int32_t | in | Kernel height. |
| kernel_w | int32_t | in | Kernel width. |
| stride_h | int32_t | in | Vertical stride. |
| stride_w | int32_t | in | Horizontal stride. |
| pad_h | int32_t | in | Top padding. |
| pad_w | int32_t | in | Left padding. |
| dilation_h | int32_t | in | Vertical dilation. |
| dilation_w | int32_t | in | Horizontal dilation. |
| out_y | int32_t | in | Output row index of the patch to pack. |
| out_x | int32_t | in | Output column index of the patch to pack. |
| pad_value | float16_t | in | Value written for taps that fall outside the input. |
| patch_row | float16_t * | out | Destination row of `kernel_h * kernel_w * in_c` elements, ordered `[kernel_h][kernel_w][in_c]`. |
Source: `Include/arm_nnsupportfunctions_flt.h:1424`
## arm_nn_softmax_1x2_f16
`function` · `c`
```c
void arm_nn_softmax_1x2_f16(const float16_t *in, float16_t *out)
```
Specialized softmax helper for a single float16 row of length 2.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in | const float16_t * | in | Pointer to two contiguous float16 input values. |
| out | float16_t * | out | Pointer to two contiguous float16 output values. |
Source: `Include/arm_nnsupportfunctions_flt.h:1447`
## arm_nn_lstm_step_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_lstm_step_f32(
const float32_t *data_in,
const float32_t *hidden_in,
float32_t *hidden_out,
const cmsis_nn_lstm_params_f32 *params,
cmsis_nn_lstm_context_f32 *buffers,
const int32_t batch_offset
)
```
Update LSTM function for an iteration step using float32 input, output and state.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const float32_t * | in | Data input pointer. |
| hidden_in | const float32_t * | in | Hidden state / recurrent input pointer. May be NULL for the first step. |
| hidden_out | float32_t * | out | Hidden state / recurrent output pointer. |
| params | const cmsis_nn_lstm_params_f32 * | in | Struct containing all information about the LSTM operator. |
| buffers | cmsis_nn_lstm_context_f32 * | in, out | Struct containing pointers to mutable cell-state storage. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params/buffers or buffers->cell_state, batch_offset <= 0). |
Source: `Include/arm_nnsupportfunctions_flt.h:1466`
## arm_nn_gru_step_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_gru_step_f32(
const float32_t *data_in,
const float32_t *hidden_in,
float32_t *hidden_out,
const cmsis_nn_gru_params_f32 *params,
cmsis_nn_gru_context_f32 *buffers,
const int32_t batch_offset
)
```
Update GRU function for a single iteration step using float32 data.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const float32_t * | in | Data input pointer for this time step. |
| hidden_in | const float32_t * | in | Recurrent input pointer. NULL for the first step (h_prev = 0). |
| hidden_out | float32_t * | out | Hidden-state output pointer for this time step. |
| params | const cmsis_nn_gru_params_f32 * | in | Struct describing the GRU operator. |
| buffers | cmsis_nn_gru_context_f32 * | in, out | Scratch buffers. temp1 (>= hidden_size) is required when reset_after == 0. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params, batch_offset <= 0, or missing temp1 when reset_after == 0). |
Source: `Include/arm_nnsupportfunctions_flt.h:1486`
---
# heliaCORE.LSTM
## arm_lstm_unidirectional_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_lstm_unidirectional_f32(
const float32_t *input,
float32_t *output,
const cmsis_nn_lstm_params_f32 *params,
cmsis_nn_lstm_context_f32 *buffers
)
```
Unidirectional LSTM inference.
:::note
On the MVE float path a NaN cell state yields a NaN hidden state, as tanh returns NaN unchanged there (#635). With cell clipping enabled the clip removes a NaN cell state first.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input sequence tensor. |
| output | float32_t * | out | Pointer to the output sequence tensor. |
| params | const cmsis_nn_lstm_params_f32 * | in | LSTM parameters and weights. |
| buffers | cmsis_nn_lstm_context_f32 * | in, out | Mutable LSTM scratch and state buffers. temp1 and temp2 are sized by `arm_lstm_unidirectional_f32_temp1_get_buffer_size()` / `arm_lstm_unidirectional_f32_temp2_get_buffer_size()`, which report 0: the float implementation never dereferences them and both may be NULL. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1781`
## arm_gru_unidirectional_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_gru_unidirectional_f32(
const float32_t *input,
float32_t *output,
const cmsis_nn_gru_params_f32 *params,
cmsis_nn_gru_context_f32 *buffers
)
```
Unidirectional GRU layer for float32 input, output and state.
Implements the reset-after GRU (Keras / TFLite default) when `params->reset_after` is non-zero, and the pre-reset variant otherwise. The hidden state is zero-initialised for the first time step, unless `buffers->hidden_state` is supplied for streaming state carry (`batch_size == 1`), in which case it seeds the initial state and receives the final hidden state on return.
:::note
NaN contract: a NaN in `input`, the previous hidden state, or the candidate gate's weight or bias reaches every output unit it feeds, on the scalar and MVE legs alike and at the shipped -Ofast: the MVE block re-establishes NaN after the table tanh with an integer-domain test that fast-math cannot elide (#251). A NaN confined to the update or reset gate's weight or bias does not reach the output: the scalar sigmoid maps NaN to 1.0 (see the note on arm_nn_sigmoid_scalar_f32 in `arm_nnsupportfunctions_flt.h`). NaN payloads and signs are not preserved on the MVE leg (default NaN, architectural). Inf follows the arithmetic.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Input sequence tensor. Must not overlap `output`: earlier outputs are re-read as the recurrent state for later time steps, so aliasing corrupts silently. |
| output | float32_t * | out | Output (hidden-state) sequence tensor. |
| params | const cmsis_nn_gru_params_f32 * | in | Struct describing the GRU operator. |
| buffers | cmsis_nn_gru_context_f32 * | in, out | Scratch buffers. May be NULL when `reset_after` != 0. temp1 is sized by `arm_gru_unidirectional_f32_temp1_get_buffer_size()`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, ARM_CMSIS_NN_ARG_ERROR otherwise. |
Source: `Include/arm_nnfunctions_flt.h:1812`
## arm_lstm_unidirectional_f32_temp1_get_buffer_size
`function` · `c`
```c
int32_t arm_lstm_unidirectional_f32_temp1_get_buffer_size(const cmsis_nn_lstm_params_f32 *lstm_params)
```
Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_f32()`.
:::note
This query reports its invalid input as -1, following the integer LSTM temp sizers (`arm_lstm_unidirectional_s8_temp1_get_buffer_size()`), not the 0 used by the float convolution and fully-connected queries in this header.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params_f32 * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f32` passed to `arm_lstm_unidirectional_f32()`. No field is read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 0 for any non-NULL lstm_params, on every build target: the float32 implementation computes its gate values per hidden unit in automatics and never dereferences temp1 or temp2, so both context pointers may be NULL. Returns -1 only for a NULL lstm_params. The query exists so arena-sizing code can treat every LSTM variant alike; a future implementation that starts staging gate vectors would change this figure, so size from the query rather than hard-coding 0. |
Source: `Include/arm_nnfunctions_flt.h:1833`
## arm_lstm_unidirectional_f32_temp2_get_buffer_size
`function` · `c`
```c
int32_t arm_lstm_unidirectional_f32_temp2_get_buffer_size(const cmsis_nn_lstm_params_f32 *lstm_params)
```
Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_f32()`. The contract is identical to `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, and the answer is the same 0 (temp2 is likewise never dereferenced).
:::note
This query reports its invalid input as -1, following the integer LSTM temp sizers (`arm_lstm_unidirectional_s8_temp1_get_buffer_size()`), not the 0 used by the float convolution and fully-connected queries in this header.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params_f32 * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f32` passed to `arm_lstm_unidirectional_f32()`. No field is read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 0 for any non-NULL lstm_params, on every build target: the float32 implementation computes its gate values per hidden unit in automatics and never dereferences temp1 or temp2, so both context pointers may be NULL. Returns -1 only for a NULL lstm_params. The query exists so arena-sizing code can treat every LSTM variant alike; a future implementation that starts staging gate vectors would change this figure, so size from the query rather than hard-coding 0. |
Source: `Include/arm_nnfunctions_flt.h:1842`
## arm_gru_unidirectional_f32_temp1_get_buffer_size
`function` · `c`
```c
int32_t arm_gru_unidirectional_f32_temp1_get_buffer_size(const cmsis_nn_gru_params_f32 *gru_params)
```
Get size of the temp1 scratch buffer required by `arm_gru_unidirectional_f32()`.
:::note
On the pre-reset path a 0 is only returned for the degenerate hidden_size == 0, which `arm_gru_unidirectional_f32()` rejects with ARM_CMSIS_NN_ARG_ERROR before any buffer access - so a 0 there never corresponds to a runnable call.
:::
:::note
This query reports an out-of-range shape as -1, following the integer LSTM temp sizers, not the 0 used by the float convolution and fully-connected queries in this header.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| gru_params | const cmsis_nn_gru_params_f32 * | in | GRU operator parameters, i.e. the same `cmsis_nn_gru_params_f32` passed to `arm_gru_unidirectional_f32()`. Only reset_after and hidden_size are read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: hidden_size * sizeof(float32_t) when reset_after == 0 (the pre-reset formulation stages the r . h_prev vector in temp1; the vector is reused across batches and time steps, so neither batch_size nor time_steps enters), and 0 when reset_after != 0 (temp1 is never dereferenced and may be NULL). Returns -1 if gru_params is NULL, if hidden_size is negative, or if the byte count would not fit in an int32_t. The figure and the range checks are the same on every build target. |
Source: `Include/arm_nnfunctions_flt.h:1863`
## arm_lstm_unidirectional_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_lstm_unidirectional_f16(
const float16_t *input,
float16_t *output,
const cmsis_nn_lstm_params_f16 *params,
cmsis_nn_lstm_context_f16 *buffers
)
```
Unidirectional LSTM inference.
:::note
On the MVE float path a NaN cell state yields a NaN hidden state, as tanh returns NaN unchanged there (#635). With cell clipping enabled the clip removes a NaN cell state first.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input sequence tensor. |
| output | float16_t * | out | Pointer to the output sequence tensor. |
| params | const cmsis_nn_lstm_params_f16 * | in | LSTM parameters and weights. |
| buffers | cmsis_nn_lstm_context_f16 * | in, out | Mutable LSTM scratch and state buffers. temp1 and temp2 are sized by `arm_lstm_unidirectional_f32_temp1_get_buffer_size()` / `arm_lstm_unidirectional_f32_temp2_get_buffer_size()`, which report 0: the float implementation never dereferences them and both may be NULL. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3558`
## arm_gru_unidirectional_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_gru_unidirectional_f16(
const float16_t *input,
float16_t *output,
const cmsis_nn_gru_params_f16 *params,
cmsis_nn_gru_context_f16 *buffers
)
```
Unidirectional GRU layer for float16 input, output and state.
Implements the reset-after GRU (Keras / TFLite default) when `params->reset_after` is non-zero, and the pre-reset variant otherwise. The hidden state is zero-initialised for the first time step, unless `buffers->hidden_state` is supplied for streaming state carry (`batch_size == 1`), in which case it seeds the initial state and receives the final hidden state on return.
:::note
NaN contract: a NaN in `input`, the previous hidden state, or the candidate gate's weight or bias reaches every output unit it feeds, on the scalar and MVE legs alike and at the shipped -Ofast: the MVE block re-establishes NaN after the table tanh with an integer-domain test that fast-math cannot elide (#251). A NaN confined to the update or reset gate's weight or bias does not reach the output: the scalar sigmoid maps NaN to 1.0 (see the note on arm_nn_sigmoid_scalar_f32 in `arm_nnsupportfunctions_flt.h`). NaN payloads and signs are not preserved on the MVE leg (default NaN, architectural). Inf follows the arithmetic.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Input sequence tensor. Must not overlap `output`: earlier outputs are re-read as the recurrent state for later time steps, so aliasing corrupts silently. |
| output | float16_t * | out | Output (hidden-state) sequence tensor. |
| params | const cmsis_nn_gru_params_f16 * | in | Struct describing the GRU operator. |
| buffers | cmsis_nn_gru_context_f16 * | in, out | Scratch buffers. May be NULL when `reset_after` != 0. temp1 is sized by `arm_gru_unidirectional_f16_temp1_get_buffer_size()`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, ARM_CMSIS_NN_ARG_ERROR otherwise. |
Source: `Include/arm_nnfunctions_flt.h:3589`
## arm_lstm_unidirectional_f16_temp1_get_buffer_size
`function` · `c`
```c
int32_t arm_lstm_unidirectional_f16_temp1_get_buffer_size(const cmsis_nn_lstm_params_f16 *lstm_params)
```
Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_f16()`. The contract is identical to `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, and the answer is the same 0 on every build target (the float16 implementation likewise never dereferences temp1 or temp2, which may both be NULL).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params_f16 * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f16` passed to `arm_lstm_unidirectional_f16()`. No field is read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 0 for any non-NULL lstm_params, -1 for a NULL lstm_params. |
Source: `Include/arm_nnfunctions_flt.h:3605`
## arm_lstm_unidirectional_f16_temp2_get_buffer_size
`function` · `c`
```c
int32_t arm_lstm_unidirectional_f16_temp2_get_buffer_size(const cmsis_nn_lstm_params_f16 *lstm_params)
```
Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_f16()`. The contract is identical to `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, and the answer is the same 0.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params_f16 * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f16` passed to `arm_lstm_unidirectional_f16()`. No field is read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 0 for any non-NULL lstm_params, -1 for a NULL lstm_params. |
Source: `Include/arm_nnfunctions_flt.h:3613`
## arm_gru_unidirectional_f16_temp1_get_buffer_size
`function` · `c`
```c
int32_t arm_gru_unidirectional_f16_temp1_get_buffer_size(const cmsis_nn_gru_params_f16 *gru_params)
```
Get size of the temp1 scratch buffer required by `arm_gru_unidirectional_f16()`. See `arm_gru_unidirectional_f32_temp1_get_buffer_size()` for the -1-on-invalid contract and the pre-reset degenerate-0 note; both apply here unchanged.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| gru_params | const cmsis_nn_gru_params_f16 * | in | GRU operator parameters, i.e. the same `cmsis_nn_gru_params_f16` passed to `arm_gru_unidirectional_f16()`. Only reset_after and hidden_size are read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: hidden_size * sizeof(float16_t) when reset_after == 0, 0 when reset_after != 0 (temp1 is never dereferenced and may be NULL). Half the figure `arm_gru_unidirectional_f32_temp1_get_buffer_size()` returns for the same shape - sizing an f16 layer with the f32 query over-allocates, and the reverse under-allocates. |
Source: `Include/arm_nnfunctions_flt.h:3628`
---
# heliaCORE.NNConv
Collection of convolution, depthwise convolution functions and their variants.
The convolution is implemented in 2 steps: im2col and General Matrix Multiplication(GEMM)
im2col is a process of converting each patch of image data into a column. After im2col, the convolution is computed as matrix-matrix multiplication.
To reduce the memory footprint, the im2col is performed partially. Each iteration, only a few column (i.e., patches) are generated followed by GEMM.
## arm_depthwise_nhwc_conv_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_nhwc_conv_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *filter_dims,
const float32_t *kernel,
const cmsis_nn_dims *bias_dims,
const float32_t *bias,
const cmsis_nn_dims *output_dims,
float32_t *output
)
```
Depthwise convolution, NHWC layout.
:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions in NHWC format. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions in NHWC-compatible depthwise format. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions in NHWC format. |
| output | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:79`
## arm_depthwise_conv_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *filter_dims,
const float32_t *kernel,
const cmsis_nn_dims *bias_dims,
const float32_t *bias,
const cmsis_nn_dims *output_dims,
float32_t *output,
arm_nn_tensor_layout layout
)
```
Depthwise convolution, dispatch by layout.
:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:119`
## arm_depthwise_conv_wrapper_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_wrapper_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *filter_dims,
const float32_t *kernel,
const cmsis_nn_dims *bias_dims,
const float32_t *bias,
const cmsis_nn_dims *output_dims,
float32_t *output
)
```
Depthwise convolution wrapper using the CMSIS-NN baseline path.
:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:158`
## arm_depthwise_conv_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_f32_get_buffer_size(
const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the temporary buffer size required by depthwise convolution.
:::note
Only one route reads scratch: on MVE builds, an NHWC depthwise with ch_mult != 1, a single input channel and at least CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD output channels can run as a regular convolution, which needs the repacked filter, `ROUND_UP(output_dims->c, 4) * filter_dims->h * filter_dims->w * sizeof(float32_t)` bytes (`ROUND_UP(output_dims->c, 8)` and `sizeof(float16_t)` for `_f16`), plus `arm_convolve_wrapper_f32_get_buffer_size` (`_f16`) for that convolution. The query reserves that for every such layer; one an exact-shape specialization takes first runs without it, so the size is an upper bound there. Every other layer the `ch_mult == 1` direct kernel and the generic kernel runs without scratch and the query returns 0 (AmbiqAI/ns-cmsis-nn#448, #625).
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:189`
## arm_depthwise_conv_wrapper_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_f32_get_buffer_size(
const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the buffer size required by the depthwise convolution wrapper.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:205`
## arm_convolve_nhwc_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_nhwc_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
Convolution, NHWC layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:232`
## arm_convolve_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data,
arm_nn_tensor_layout layout
)
```
Convolution, dispatch by layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:266`
## arm_convolve_wrapper_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_wrapper_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
Convolution wrapper using the CMSIS-NN baseline path.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:294`
## arm_convolve_1x1_nhwc_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_nhwc_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
1x1 convolution, NHWC layout.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:321`
## arm_convolve_1x1_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data,
arm_nn_tensor_layout layout
)
```
1x1 convolution, dispatch by layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:354`
## arm_convolve_1_x_n_nhwc_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
1xN convolution, NHWC layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:391`
## arm_convolve_1_x_n_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1_x_n_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data,
arm_nn_tensor_layout layout
)
```
1xN convolution, dispatch by layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:428`
## arm_convolve_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_f32_get_buffer_size(
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the temporary buffer size required by convolution.
:::note
When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still reports only the temporary input/im2col scratch requirement. Any offline-packed filter storage is expected to be provided by the caller.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:456`
## arm_convolve_wrapper_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_wrapper_f32_get_buffer_size(
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the buffer size required by the convolution wrapper.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:472`
## arm_convolve_1x1_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_1x1_f32_get_buffer_size(
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the buffer size required by 1x1 convolution.
:::note
Returns `0` for the unity-stride no-pack path. For non-unity-stride NHWC 1x1 convolution, the returned scratch size enables the packed-tile + GEMM path. When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still excludes the offline-packed filter storage itself.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:493`
## arm_convolve_1_x_n_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_1_x_n_f32_get_buffer_size(
const cmsis_nn_conv_params_f32 *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the buffer size required by 1xN convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:510`
## arm_transpose_conv_wrapper_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_conv_wrapper_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data,
arm_nn_tensor_layout layout
)
```
Transpose convolution wrapper using the CMSIS-NN baseline path.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1540`
## arm_transpose_conv_nhwc_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_conv_nhwc_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
Transpose convolution, NHWC layout.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1570`
## arm_transpose_conv_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_conv_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data,
arm_nn_tensor_layout layout
)
```
Transpose convolution, dispatch by layout.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1600`
## arm_transpose_conv_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_transpose_conv_f32_get_buffer_size(
const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *out_dims
)
```
Get the temporary buffer size required by transpose convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| out_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:1623`
## arm_transpose_conv_f32_get_reverse_conv_buffer_size
`function` · `c`
```c
int32_t arm_transpose_conv_f32_get_reverse_conv_buffer_size(
const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims
)
```
Get the reverse-convolution workspace size used by transpose convolution helpers.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no reverse-convolution buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:1638`
## arm_depthwise_nhwc_conv_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_nhwc_conv_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output
)
```
Depthwise convolution, NHWC layout.
:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16` unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions in NHWC format. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions in NHWC-compatible depthwise format. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions in NHWC format. |
| output | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2224`
## arm_depthwise_nhwc_conv_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_nhwc_conv_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output
)
```
Depthwise convolution, NHWC layout.
:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16` unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.
:::
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_nhwc_conv_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_nhwc_conv_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions in NHWC format. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions in NHWC-compatible depthwise format. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions in NHWC format. |
| output | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2243`
## arm_depthwise_conv_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output,
arm_nn_tensor_layout layout
)
```
Depthwise convolution, dispatch by layout.
:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16` unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2273`
## arm_depthwise_conv_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output,
arm_nn_tensor_layout layout
)
```
Depthwise convolution, dispatch by layout.
:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16` unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.
:::
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_conv_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_conv_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2293`
## arm_depthwise_conv_wrapper_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_wrapper_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output
)
```
Depthwise convolution wrapper using the CMSIS-NN baseline path.
:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16` unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2324`
## arm_depthwise_conv_wrapper_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_wrapper_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output
)
```
Depthwise convolution wrapper using the CMSIS-NN baseline path.
:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.
:::
:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16` unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.
:::
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_conv_wrapper_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_conv_wrapper_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2343`
## arm_depthwise_conv_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_f16_get_buffer_size(
const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the temporary buffer size required by depthwise convolution.
:::note
Only one route reads scratch: on MVE builds, an NHWC depthwise with ch_mult != 1, a single input channel and at least CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD output channels can run as a regular convolution, which needs the repacked filter, `ROUND_UP(output_dims->c, 4) * filter_dims->h * filter_dims->w * sizeof(float32_t)` bytes (`ROUND_UP(output_dims->c, 8)` and `sizeof(float16_t)` for `_f16`), plus `arm_convolve_wrapper_f32_get_buffer_size` (`_f16`) for that convolution. The query reserves that for every such layer; one an exact-shape specialization takes first runs without it, so the size is an upper bound there. Every other layer the `ch_mult == 1` direct kernel and the generic kernel runs without scratch and the query returns 0 (AmbiqAI/ns-cmsis-nn#448, #625).
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:2357`
## arm_depthwise_conv_wrapper_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_f16_get_buffer_size(
const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the buffer size required by the depthwise convolution wrapper.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:2366`
## arm_convolve_nhwc_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_nhwc_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
Convolution, NHWC layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2374`
## arm_convolve_nhwc_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_nhwc_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
Convolution, NHWC layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.
:::
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_nhwc_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2393`
## arm_convolve_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
arm_nn_tensor_layout layout
)
```
Convolution, dispatch by layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.
:::
:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct OHWI / NT_N_PACKED fallback accumulates bias and every tap in float32 and rounds to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #457); the 1x1, 1xN and patch-GEMM paths go through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16, whose scalar legs do the same, as do the 1xN no-padding OHWI region and the k=3 / k=5 conv1d specializations (#465). MVE leg: the direct small-C kernel accumulates in float32 (widened lanes); the direct OHWI / NT_N_PACKED fallback, every matmul-backed path (1x1, 1xN, patch-GEMM), the 1xN no-padding region and the conv1d specializations use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. Where a dot product spreads its taps over the lanes of one vector (OHWI rows, the contiguous-K matmul, the conv1d k=3 / k=5 OHWI kernels), each lane's own taps form its blocks and, once the reduction exceeds 32 taps, each block's lanes are widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (the first block sets it); the four pair accumulators are summed once as (0+1) + (2+3), the bias is added in float32 and the total rounds once. The k=3 / k=5 kernels close a block on a whole input-channel step (30 taps per lane). An output's taps are the ones its kernel multiplies: the direct fallback skips padded taps, so an edge output counts only its in-range taps, while patch-GEMM and the 1xN padded regions multiply a zero-padded patch and count its padded taps too. Patch-GEMM runs only when ctx provides its scratch, so an edge output's value can depend on whether ctx->buf is given. The fold's order is fixed; the float16 reduction of a dot of at most 32 taps is left to the compiler, which may reorder it under -ffast-math, as before #586.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2431`
## arm_convolve_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
arm_nn_tensor_layout layout
)
```
Convolution, dispatch by layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.
:::
:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct OHWI / NT_N_PACKED fallback accumulates bias and every tap in float32 and rounds to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #457); the 1x1, 1xN and patch-GEMM paths go through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16, whose scalar legs do the same, as do the 1xN no-padding OHWI region and the k=3 / k=5 conv1d specializations (#465). MVE leg: the direct small-C kernel accumulates in float32 (widened lanes); the direct OHWI / NT_N_PACKED fallback, every matmul-backed path (1x1, 1xN, patch-GEMM), the 1xN no-padding region and the conv1d specializations use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. Where a dot product spreads its taps over the lanes of one vector (OHWI rows, the contiguous-K matmul, the conv1d k=3 / k=5 OHWI kernels), each lane's own taps form its blocks and, once the reduction exceeds 32 taps, each block's lanes are widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (the first block sets it); the four pair accumulators are summed once as (0+1) + (2+3), the bias is added in float32 and the total rounds once. The k=3 / k=5 kernels close a block on a whole input-channel step (30 taps per lane). An output's taps are the ones its kernel multiplies: the direct fallback skips padded taps, so an edge output counts only its in-range taps, while patch-GEMM and the 1xN padded regions multiply a zero-padded patch and count its padded taps too. Patch-GEMM runs only when ctx provides its scratch, so an edge output's value can depend on whether ctx->buf is given. The fold's order is fixed; the float16 reduction of a dot of at most 32 taps is left to the compiler, which may reorder it under -ffast-math, as before #586.
:::
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2451`
## arm_convolve_wrapper_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_wrapper_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
Convolution wrapper using the CMSIS-NN baseline path.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2466`
## arm_convolve_wrapper_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_wrapper_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
Convolution wrapper using the CMSIS-NN baseline path.
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_wrapper_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_wrapper_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2485`
## arm_convolve_1x1_nhwc_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_nhwc_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
1x1 convolution, NHWC layout.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2499`
## arm_convolve_1x1_nhwc_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_nhwc_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
1x1 convolution, NHWC layout.
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1x1_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1x1_nhwc_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2518`
## arm_convolve_1x1_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
arm_nn_tensor_layout layout
)
```
1x1 convolution, dispatch by layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2532`
## arm_convolve_1x1_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
arm_nn_tensor_layout layout
)
```
1x1 convolution, dispatch by layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.
:::
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1x1_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1x1_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2552`
## arm_convolve_1_x_n_nhwc_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
1xN convolution, NHWC layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2567`
## arm_convolve_1_x_n_nhwc_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
1xN convolution, NHWC layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.
:::
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1_x_n_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1_x_n_nhwc_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2586`
## arm_convolve_1_x_n_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1_x_n_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
arm_nn_tensor_layout layout
)
```
1xN convolution, dispatch by layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.
:::
:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and every product accumulate in float32 and round to float16 once (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: the padded regions go through the matmul helpers and the no-padding region through a strided kernel, all with blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. The padded regions multiply a zero-padded patch row, so their outputs count the padded taps as well.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2610`
## arm_convolve_1_x_n_f16_acc16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1_x_n_f16_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
arm_nn_tensor_layout layout
)
```
1xN convolution, dispatch by layout.
:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.
:::
:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and every product accumulate in float32 and round to float16 once (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: the padded regions go through the matmul helpers and the no-padding region through a strided kernel, all with blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. The padded regions multiply a zero-padded patch row, so their outputs count the padded taps as well.
:::
:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1_x_n_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1_x_n_f16.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2630`
## arm_convolve_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_f16_get_buffer_size(
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the temporary buffer size required by convolution.
:::note
When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still reports only the temporary input/im2col scratch requirement. Any offline-packed filter storage is expected to be provided by the caller.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:2645`
## arm_convolve_wrapper_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_wrapper_f16_get_buffer_size(
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the buffer size required by the convolution wrapper.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:2654`
## arm_convolve_1x1_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_1x1_f16_get_buffer_size(
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the buffer size required by 1x1 convolution.
:::note
Returns `0` for the unity-stride no-pack path. For non-unity-stride NHWC 1x1 convolution, the returned scratch size enables the packed-tile + GEMM path. When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still excludes the offline-packed filter storage itself.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:2662`
## arm_convolve_1_x_n_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_1_x_n_f16_get_buffer_size(
const cmsis_nn_conv_params_f16 *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
arm_nn_tensor_layout layout
)
```
Get the buffer size required by 1xN convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:2671`
## arm_transpose_conv_wrapper_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_conv_wrapper_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
arm_nn_tensor_layout layout
)
```
Transpose convolution wrapper using the CMSIS-NN baseline path.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3354`
## arm_transpose_conv_nhwc_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_conv_nhwc_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
Transpose convolution, NHWC layout.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3370`
## arm_transpose_conv_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_conv_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
arm_nn_tensor_layout layout
)
```
Transpose convolution, dispatch by layout.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3385`
## arm_transpose_conv_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_transpose_conv_f16_get_buffer_size(
const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *out_dims
)
```
Get the temporary buffer size required by transpose convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| out_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no scratch buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:3401`
## arm_transpose_conv_f16_get_reverse_conv_buffer_size
`function` · `c`
```c
int32_t arm_transpose_conv_f16_get_reverse_conv_buffer_size(
const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims
)
```
Get the reverse-convolution workspace size used by transpose convolution helpers.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes, or 0 when no reverse-convolution buffer is needed. |
Source: `Include/arm_nnfunctions_flt.h:3410`
---
# heliaCORE.NNSupport
## arm_transpose_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_transpose_params_f32 *params,
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *output_dims,
float32_t *output
)
```
Transpose a floating-point tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| params | const cmsis_nn_transpose_params_f32 * | in | Transpose parameters, including permutation and layout information. num_dims must be in [1, 4] and perm must be a bijection over [0, num_dims - 1]. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. The first params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[perm[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not. |
| output | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1135`
## arm_concatenation_f32_x
`function` · `c`
```c
void arm_concatenation_f32_x(
const float32_t *input,
int32_t input_x,
int32_t input_y,
int32_t input_z,
int32_t input_w,
float32_t *output,
int32_t output_x,
uint32_t offset_x
)
```
Concatenate tensors along the X axis.
Call once per input tensor: `offset_x` selects where the input is stored along the X axis of the output tensor and must be advanced by `input_x` after each call. The output tensor must have the same height, channels and batch size as every input tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float32_t * | out | Pointer to the output tensor. |
| output_x | int32_t | in | Width of the output tensor. |
| offset_x | uint32_t | in | Offset on the X axis at which the input tensor is stored. Must be less than `output_x`. |
Source: `Include/arm_nnfunctions_flt.h:1158`
## arm_concatenation_f32_y
`function` · `c`
```c
void arm_concatenation_f32_y(
const float32_t *input,
int32_t input_x,
int32_t input_y,
int32_t input_z,
int32_t input_w,
float32_t *output,
int32_t output_y,
uint32_t offset_y
)
```
Concatenate tensors along the Y axis.
Call once per input tensor: `offset_y` selects where the input is stored along the Y axis of the output tensor and must be advanced by `input_y` after each call. The output tensor must have the same width, channels and batch size as every input tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float32_t * | out | Pointer to the output tensor. |
| output_y | int32_t | in | Height of the output tensor. |
| offset_y | uint32_t | in | Offset on the Y axis at which the input tensor is stored. Must be less than `output_y`. |
Source: `Include/arm_nnfunctions_flt.h:1183`
## arm_concatenation_f32_z
`function` · `c`
```c
void arm_concatenation_f32_z(
const float32_t *input,
int32_t input_x,
int32_t input_y,
int32_t input_z,
int32_t input_w,
float32_t *output,
int32_t output_z,
uint32_t offset_z
)
```
Concatenate tensors along the Z axis.
Call once per input tensor: `offset_z` selects where the input is stored along the Z axis of the output tensor and must be advanced by `input_z` after each call. The output tensor must have the same width, height and batch size as every input tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float32_t * | out | Pointer to the output tensor. |
| output_z | int32_t | in | Channels in the output tensor. |
| offset_z | uint32_t | in | Offset on the Z axis at which the input tensor is stored. Must be less than `output_z`. |
Source: `Include/arm_nnfunctions_flt.h:1208`
## arm_concatenation_f32_w
`function` · `c`
```c
void arm_concatenation_f32_w(
const float32_t *input,
int32_t input_x,
int32_t input_y,
int32_t input_z,
int32_t input_w,
float32_t *output,
uint32_t offset_w
)
```
Concatenate tensors along the W axis.
Call once per input tensor: `offset_w` selects where the input is stored along the W axis of the output tensor and must be advanced by `input_w` after each call. The output tensor must have the same width, height and channels as every input tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float32_t * | out | Pointer to the output tensor. |
| offset_w | uint32_t | in | Offset on the W axis at which the input tensor is stored. |
Source: `Include/arm_nnfunctions_flt.h:1232`
## arm_concatenation_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_concatenation_f32(
const float32_t *const *input_data,
int32_t num_inputs,
const int32_t *axis_sizes,
int32_t output_dims,
const int32_t *output_shape,
int32_t axis,
float32_t *output_data
)
```
Concatenate float32 tensors of any rank along one axis.
Rank-agnostic sibling of the 4-D per-axis arm_concatenation_f32_{x,y,z,w} entry points: all inputs at once, any rank, any axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Input `s` has the output shape with `output_shape`[axis] replaced by `axis_sizes`[s]; the inputs are laid down in order along the axis. Inputs must not overlap the output. A dimension of 0 is accepted and copies nothing.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t *const * | in | Array of `num_inputs` pointers to the flattened (row-major) inputs. |
| num_inputs | int32_t | in | Number of inputs (>= 1). |
| axis_sizes | const int32_t * | in | Array of length `num_inputs:` each input's extent along `axis`. |
| output_dims | int32_t | in | Number of dimensions in `output_shape` (>= 1). |
| output_shape | const int32_t * | in | Output shape; `output_shape`[axis] must equal the sum of `axis_sizes`. |
| axis | int32_t | in | Axis to concatenate along (0 <= axis < output_dims). |
| output_data | float32_t * | out | Pointer to the flattened output. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, size entry, size sum, NULL pointer or an element count above INT32_MAX. |
Source: `Include/arm_nnfunctions_flt.h:1259`
## arm_split_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_split_f32(
const float32_t *input_data,
int32_t input_dims,
const int32_t *input_shape,
int32_t axis,
int32_t num_splits,
const int32_t *split_dims,
float32_t *const *output_data
)
```
Split a float32 tensor of any rank into several tensors along one axis.
Inverse of arm_concatenation_f32; per-split lengths also cover SPLIT_V. Output `s` has the input shape with `input_shape`[axis] replaced by `split_dims`[s]. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input. A dimension of 0 is accepted and copies nothing.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Pointer to the flattened (row-major) input. |
| input_dims | int32_t | in | Number of dimensions in `input_shape` (>= 1). |
| input_shape | const int32_t * | in | Input shape; `input_shape`[axis] must equal the sum of `split_dims`. |
| axis | int32_t | in | Axis to split along (0 <= axis < input_dims). |
| num_splits | int32_t | in | Number of outputs (>= 1). |
| split_dims | const int32_t * | in | Array of length `num_splits:` each output's extent along `axis`. |
| output_data | float32_t *const * | out | Array of `num_splits` pointers to the flattened outputs. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, split entry, split sum, NULL pointer or an element count above INT32_MAX. |
Source: `Include/arm_nnfunctions_flt.h:1285`
## arm_pack_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_pack_f32(
const float32_t *const *input_data,
int32_t num_inputs,
int32_t input_dims,
const int32_t *input_shape,
int32_t axis,
float32_t *output_data
)
```
Stack float32 tensors of equal shape along a new axis (TFLite PACK).
The output shape is `input_shape` with `num_inputs` inserted at `axis`; input `s` lands at index `s` of that axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Inputs must not overlap the output. Rank-0 inputs (`input_dims` == 0, `axis` == 0) stack into a vector.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t *const * | in | Array of `num_inputs` pointers to the flattened (row-major) inputs. |
| num_inputs | int32_t | in | Number of inputs (>= 1). |
| input_dims | int32_t | in | Number of dimensions of each input (>= 0). |
| input_shape | const int32_t * | in | Shape shared by every input (may be NULL when `input_dims` is 0). |
| axis | int32_t | in | Position of the new axis in the output (0 <= axis <= input_dims). |
| output_data | float32_t * | out | Pointer to the flattened output. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, NULL pointer or an element count above INT32_MAX. |
Source: `Include/arm_nnfunctions_flt.h:1310`
## arm_unpack_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_unpack_f32(
const float32_t *input_data,
int32_t input_dims,
const int32_t *input_shape,
int32_t axis,
float32_t *const *output_data
)
```
Unstack a float32 tensor along one axis into `input_shape`[axis] tensors (TFLite UNPACK).
Inverse of arm_pack_f32: output `s` is the input with the axis fixed at index `s` and removed from the shape. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Pointer to the flattened (row-major) input. |
| input_dims | int32_t | in | Number of dimensions in `input_shape` (>= 1). |
| input_shape | const int32_t * | in | Input shape; `input_shape`[axis] (>= 1) is the number of outputs. |
| axis | int32_t | in | Axis to unstack (0 <= axis < input_dims). |
| output_data | float32_t *const * | out | Array of `input_shape`[axis] pointers to the flattened outputs. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, a zero-extent unstack axis (no outputs to produce), NULL pointer or an element count above INT32_MAX. |
Source: `Include/arm_nnfunctions_flt.h:1334`
## arm_batch_norm_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_batch_norm_f32(
const float32_t *input,
float32_t *output,
const float32_t *scale,
const float32_t *bias,
const cmsis_nn_dims *input_dims,
arm_nn_tensor_layout layout
)
```
Apply batch normalization.
Computes `output = input * scale[c] + bias[c]` for every element of channel `c`, with `scale` and `bias` holding the pre-folded per-channel factors.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor data. Format: [N, H, W, C]. |
| output | float32_t * | out | Pointer to the output tensor data, same shape as `input`. |
| scale | const float32_t * | in | Per-channel scale, `input_dims->c` values. |
| bias | const float32_t * | in | Per-channel bias, `input_dims->c` values. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Every dimension must be positive. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Must be `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1390`
## arm_reshape_f32
`function` · `c`
```c
void arm_reshape_f32(const float32_t *input, float32_t *output, uint32_t total_size)
```
Reshape by copying data without changing element order.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor data. |
| output | float32_t * | out | Pointer to the output tensor data. Nothing is copied when it aliases `input`. |
| total_size | uint32_t | in | Number of elements to copy. |
Source: `Include/arm_nnfunctions_flt.h:1404`
## arm_transpose_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_transpose_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *output_dims,
float16_t *output
)
```
Transpose a floating-point tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| params | const cmsis_nn_transpose_params_f16 * | in | Transpose parameters, including permutation and layout information. num_dims must be in [1, 4] and perm must be a bijection over [0, num_dims - 1]. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. The first params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[perm[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not. |
| output | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3161`
## arm_concatenation_f16_x
`function` · `c`
```c
void arm_concatenation_f16_x(
const float16_t *input,
int32_t input_x,
int32_t input_y,
int32_t input_z,
int32_t input_w,
float16_t *output,
int32_t output_x,
uint32_t offset_x
)
```
Concatenate tensors along the X axis.
Call once per input tensor: `offset_x` selects where the input is stored along the X axis of the output tensor and must be advanced by `input_x` after each call. The output tensor must have the same height, channels and batch size as every input tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float16_t * | out | Pointer to the output tensor. |
| output_x | int32_t | in | Width of the output tensor. |
| offset_x | uint32_t | in | Offset on the X axis at which the input tensor is stored. Must be less than `output_x`. |
Source: `Include/arm_nnfunctions_flt.h:3171`
## arm_concatenation_f16_y
`function` · `c`
```c
void arm_concatenation_f16_y(
const float16_t *input,
int32_t input_x,
int32_t input_y,
int32_t input_z,
int32_t input_w,
float16_t *output,
int32_t output_y,
uint32_t offset_y
)
```
Concatenate tensors along the Y axis.
Call once per input tensor: `offset_y` selects where the input is stored along the Y axis of the output tensor and must be advanced by `input_y` after each call. The output tensor must have the same width, channels and batch size as every input tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float16_t * | out | Pointer to the output tensor. |
| output_y | int32_t | in | Height of the output tensor. |
| offset_y | uint32_t | in | Offset on the Y axis at which the input tensor is stored. Must be less than `output_y`. |
Source: `Include/arm_nnfunctions_flt.h:3183`
## arm_concatenation_f16_z
`function` · `c`
```c
void arm_concatenation_f16_z(
const float16_t *input,
int32_t input_x,
int32_t input_y,
int32_t input_z,
int32_t input_w,
float16_t *output,
int32_t output_z,
uint32_t offset_z
)
```
Concatenate tensors along the Z axis.
Call once per input tensor: `offset_z` selects where the input is stored along the Z axis of the output tensor and must be advanced by `input_z` after each call. The output tensor must have the same width, height and batch size as every input tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float16_t * | out | Pointer to the output tensor. |
| output_z | int32_t | in | Channels in the output tensor. |
| offset_z | uint32_t | in | Offset on the Z axis at which the input tensor is stored. Must be less than `output_z`. |
Source: `Include/arm_nnfunctions_flt.h:3195`
## arm_concatenation_f16_w
`function` · `c`
```c
void arm_concatenation_f16_w(
const float16_t *input,
int32_t input_x,
int32_t input_y,
int32_t input_z,
int32_t input_w,
float16_t *output,
uint32_t offset_w
)
```
Concatenate tensors along the W axis.
Call once per input tensor: `offset_w` selects where the input is stored along the W axis of the output tensor and must be advanced by `input_w` after each call. The output tensor must have the same width, height and channels as every input tensor.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float16_t * | out | Pointer to the output tensor. |
| offset_w | uint32_t | in | Offset on the W axis at which the input tensor is stored. |
Source: `Include/arm_nnfunctions_flt.h:3207`
## arm_concatenation_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_concatenation_f16(
const float16_t *const *input_data,
int32_t num_inputs,
const int32_t *axis_sizes,
int32_t output_dims,
const int32_t *output_shape,
int32_t axis,
float16_t *output_data
)
```
Concatenate float32 tensors of any rank along one axis.
Rank-agnostic sibling of the 4-D per-axis arm_concatenation_f32_{x,y,z,w} entry points: all inputs at once, any rank, any axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Input `s` has the output shape with `output_shape`[axis] replaced by `axis_sizes`[s]; the inputs are laid down in order along the axis. Inputs must not overlap the output. A dimension of 0 is accepted and copies nothing.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t *const * | in | Array of `num_inputs` pointers to the flattened (row-major) inputs. |
| num_inputs | int32_t | in | Number of inputs (>= 1). |
| axis_sizes | const int32_t * | in | Array of length `num_inputs:` each input's extent along `axis`. |
| output_dims | int32_t | in | Number of dimensions in `output_shape` (>= 1). |
| output_shape | const int32_t * | in | Output shape; `output_shape`[axis] must equal the sum of `axis_sizes`. |
| axis | int32_t | in | Axis to concatenate along (0 <= axis < output_dims). |
| output_data | float16_t * | out | Pointer to the flattened output. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, size entry, size sum, NULL pointer or an element count above INT32_MAX. |
Source: `Include/arm_nnfunctions_flt.h:3218`
## arm_pack_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_pack_f16(
const float16_t *const *input_data,
int32_t num_inputs,
int32_t input_dims,
const int32_t *input_shape,
int32_t axis,
float16_t *output_data
)
```
Stack float32 tensors of equal shape along a new axis (TFLite PACK).
The output shape is `input_shape` with `num_inputs` inserted at `axis`; input `s` lands at index `s` of that axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Inputs must not overlap the output. Rank-0 inputs (`input_dims` == 0, `axis` == 0) stack into a vector.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t *const * | in | Array of `num_inputs` pointers to the flattened (row-major) inputs. |
| num_inputs | int32_t | in | Number of inputs (>= 1). |
| input_dims | int32_t | in | Number of dimensions of each input (>= 0). |
| input_shape | const int32_t * | in | Shape shared by every input (may be NULL when `input_dims` is 0). |
| axis | int32_t | in | Position of the new axis in the output (0 <= axis <= input_dims). |
| output_data | float16_t * | out | Pointer to the flattened output. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, NULL pointer or an element count above INT32_MAX. |
Source: `Include/arm_nnfunctions_flt.h:3229`
## arm_unpack_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_unpack_f16(
const float16_t *input_data,
int32_t input_dims,
const int32_t *input_shape,
int32_t axis,
float16_t *const *output_data
)
```
Unstack a float32 tensor along one axis into `input_shape`[axis] tensors (TFLite UNPACK).
Inverse of arm_pack_f32: output `s` is the input with the axis fixed at index `s` and removed from the shape. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to the flattened (row-major) input. |
| input_dims | int32_t | in | Number of dimensions in `input_shape` (>= 1). |
| input_shape | const int32_t * | in | Input shape; `input_shape`[axis] (>= 1) is the number of outputs. |
| axis | int32_t | in | Axis to unstack (0 <= axis < input_dims). |
| output_data | float16_t *const * | out | Array of `input_shape`[axis] pointers to the flattened outputs. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, a zero-extent unstack axis (no outputs to produce), NULL pointer or an element count above INT32_MAX. |
Source: `Include/arm_nnfunctions_flt.h:3239`
## arm_batch_norm_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_batch_norm_f16(
const float16_t *input,
float16_t *output,
const float16_t *scale,
const float16_t *bias,
const cmsis_nn_dims *input_dims,
arm_nn_tensor_layout layout
)
```
Apply batch normalization.
Computes `output = input * scale[c] + bias[c]` for every element of channel `c`, with `scale` and `bias` holding the pre-folded per-channel factors.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor data. Format: [N, H, W, C]. |
| output | float16_t * | out | Pointer to the output tensor data, same shape as `input`. |
| scale | const float16_t * | in | Per-channel scale, `input_dims->c` values. |
| bias | const float16_t * | in | Per-channel bias, `input_dims->c` values. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Every dimension must be positive. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Must be `ARM_NN_LAYOUT_NHWC`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3272`
## arm_reshape_f16
`function` · `c`
```c
void arm_reshape_f16(const float16_t *input, float16_t *output, uint32_t total_size)
```
Reshape by copying data without changing element order.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor data. |
| output | float16_t * | out | Pointer to the output tensor data. Nothing is copied when it aliases `input`. |
| total_size | uint32_t | in | Number of elements to copy. |
Source: `Include/arm_nnfunctions_flt.h:3282`
---
# heliaCORE.Pad
## arm_pad_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_pad_f32(
const float32_t *input,
float32_t *output,
float32_t pad_value,
const cmsis_nn_dims *input_size,
const cmsis_nn_dims *pre_pad,
const cmsis_nn_dims *post_pad
)
```
Pad a tensor with a constant value.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor data. |
| output | float32_t * | out | Pointer to the output tensor data, sized by `input_size` plus `pre_pad` and `post_pad` in every dimension. |
| pad_value | float32_t | in | Value to pad with. |
| input_size | const cmsis_nn_dims * | in | Input tensor dimensions. |
| pre_pad | const cmsis_nn_dims * | in | Padding to apply before the data in each dimension. |
| post_pad | const cmsis_nn_dims * | in | Padding to apply after the data in each dimension. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL or a padded output dimension is not positive. |
Source: `Include/arm_nnfunctions_flt.h:1361`
## arm_pad_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_pad_f16(
const float16_t *input,
float16_t *output,
float16_t pad_value,
const cmsis_nn_dims *input_size,
const cmsis_nn_dims *pre_pad,
const cmsis_nn_dims *post_pad
)
```
Pad a tensor with a constant value.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor data. |
| output | float16_t * | out | Pointer to the output tensor data, sized by `input_size` plus `pre_pad` and `post_pad` in every dimension. |
| pad_value | float16_t | in | Value to pad with. |
| input_size | const cmsis_nn_dims * | in | Input tensor dimensions. |
| pre_pad | const cmsis_nn_dims * | in | Padding to apply before the data in each dimension. |
| post_pad | const cmsis_nn_dims * | in | Padding to apply after the data in each dimension. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL or a padded output dimension is not positive. |
Source: `Include/arm_nnfunctions_flt.h:3255`
---
# heliaCORE.Pooling
Perform max and average pooling operations
## arm_max_pool_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_max_pool_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params_f32 *pool_params,
const cmsis_nn_dims *input_dims,
const float32_t *src,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
float32_t *dst
)
```
Max pooling.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| pool_params | const cmsis_nn_pool_params_f32 * | in | Pooling parameters (stride, padding and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| src | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Pooling kernel dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| dst | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success, including an output with no rows or no columns (an extent of 0 or less), which writes nothing; `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments: a NULL pointer argument other than ctx, a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then. |
Source: `Include/arm_nnfunctions_flt.h:540`
## arm_avg_pool_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_avg_pool_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params_f32 *pool_params,
const cmsis_nn_dims *input_dims,
const float32_t *src,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
float32_t *dst
)
```
Average pooling.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| pool_params | const cmsis_nn_pool_params_f32 * | in | Pooling parameters (stride, padding and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| src | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Pooling kernel dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| dst | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success, including an output with no rows or no columns (an extent of 0 or less), which writes nothing; `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments: a NULL pointer argument other than ctx, a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then. |
Source: `Include/arm_nnfunctions_flt.h:565`
## arm_max_pool_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_max_pool_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params_f16 *pool_params,
const cmsis_nn_dims *input_dims,
const float16_t *src,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
float16_t *dst
)
```
Max pooling.
:::note
The output activation clamp on the scalar (non-MVE) build path is the bit-classified clamp of #380, so a NaN that reaches the clamp comes back as NaN at every optimization level on the gated toolchains rather than as a bound. A NaN rarely reaches it, though: the scalar max reduction uses an ordered compare that drops a NaN window element (and its NaN behavior at the shipped -Ofast is unspecified), and the MVE path's vmaxnmq reduction and vmaxnmq/vminnmq clamp suppress NaN, so this kernel does not promise NaN propagation end to end.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| pool_params | const cmsis_nn_pool_params_f16 * | in | Pooling parameters (stride, padding and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| src | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Pooling kernel dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| dst | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success, including an output with no rows or no columns (an extent of 0 or less), which writes nothing; `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments: a NULL pointer argument other than ctx, a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then. |
Source: `Include/arm_nnfunctions_flt.h:2695`
## arm_avg_pool_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_avg_pool_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params_f16 *pool_params,
const cmsis_nn_dims *input_dims,
const float16_t *src,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
float16_t *dst
)
```
Average pooling.
:::note
On non-MVE builds every output element goes through the bit-classified scalar clamp of #380, so a NaN in the pooling window propagates through the window sum and the output activation clamp to the output element at every optimization level on the gated toolchains, including the shipped -Ofast. On MVE builds the clamp is vmaxnmq/vminnmq with no NaN restore, so a NaN resolves to a clamp bound there instead.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| pool_params | const cmsis_nn_pool_params_f16 * | in | Pooling parameters (stride, padding and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| src | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Pooling kernel dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| dst | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success, including an output with no rows or no columns (an extent of 0 or less), which writes nothing; `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments: a NULL pointer argument other than ctx, a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then. |
Source: `Include/arm_nnfunctions_flt.h:2712`
---
# heliaCORE.Public
A collection of functions to perform basic operations for neural network layers. Functions with a _s8 suffix support TensorFlow Lite framework.
---
# heliaCORE.Quantization
## arm_dequantize_f16_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_dequantize_f16_f32(const float16_t *input, float32_t *output, int32_t block_size)
```
Widen a float16 vector to float32.
Bit-exact widening of every input class: finite values, subnormals (normal in float32), +/-0 and +/-Inf convert exactly. No accumulation, no rounding. NaN behavior: on every leg a NaN stays a NaN with its sign, quiet bit and payload preserved bit-exactly (a signaling NaN stays signaling). The scalar leg widens on integer lanes and raises no floating-point exception flag. The MVE leg converts each 8-element block with the vector VCVT first and then rebuilds the NaN lanes from the half's bits (per 4-lane vector, 8 elements per main-loop block), so a signaling-NaN input may leave FPSCR.IOC (invalid operation, cumulative) set on that leg; no trap, and the result is the same bits. Input and output must not overlap. Serves the f16-weights DEQUANTIZE op (`kws_float_fp16_weights`).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the float16 input vector. |
| output | float32_t * | out | Pointer to the float32 output vector. |
| block_size | int32_t | in | Number of elements (0 is a no-op). |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or a pointer is NULL with a non-zero `block_size`. |
Source: `Include/arm_nnfunctions_flt.h:2918`
---
# heliaCORE.Reduction
## arm_reduce_sum_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_sum_f32(
const float32_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Computes the sum of the input tensor along the specified axes.
Sums are accumulated in float32 (also for the float16 variant, which rounds once to float16 at the end), so results do not overflow at float16 range and precision does not degrade with the reduction count. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (4D NHWC) |
| axis_dims | const cmsis_nn_dims * | in | 4D binary axis mask (non-zero = reduce that axis) |
| output_data | float32_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (reduced axes have size 1) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1902`
## arm_argmin_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_argmin_f32(
const float32_t *input_data,
const cmsis_nn_dims *input_dims,
int32_t axis,
int32_t *output_data
)
```
Returns the first minimum's axis-relative INT32 index for a f32 tensor.
The input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.
A NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.
Metadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis | int32_t | in | Canonical reduction axis, in [0,3]. |
| output_data | int32_t * | out | INT32 indices, each in [0,input_dims[axis]). |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:1940`
## arm_argmax_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_argmax_f32(
const float32_t *input_data,
const cmsis_nn_dims *input_dims,
int32_t axis,
int32_t *output_data
)
```
Returns the first maximum's axis-relative INT32 index for a f32 tensor.
The input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.
A NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.
Metadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis | int32_t | in | Canonical reduction axis, in [0,3]. |
| output_data | int32_t * | out | INT32 indices, each in [0,input_dims[axis]). |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:1974`
## arm_reduce_max_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_max_f32(
const float32_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Reduces a f32 NHWC tensor to its maximum along a binary axis mask.
Values are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7fc00000); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.
A zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces -Inf; an empty output performs no accesses to data buffers.
All metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis_dims | const cmsis_nn_dims * | in | Four binary reduction flags. |
| output_data | float32_t * | out | Output tensor. |
| output_dims | const cmsis_nn_dims * | in | NHWC output shape with reduced axes retained as 1. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write. |
Source: `Include/arm_nnfunctions_flt.h:2004`
## arm_reduce_min_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_min_f32(
const float32_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Reduces a f32 NHWC tensor to its minimum along a binary axis mask.
Values are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7fc00000); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.
A zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces +Inf; an empty output performs no accesses to data buffers.
All metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis_dims | const cmsis_nn_dims * | in | Four binary reduction flags. |
| output_data | float32_t * | out | Output tensor. |
| output_dims | const cmsis_nn_dims * | in | NHWC output shape with reduced axes retained as 1. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write. |
Source: `Include/arm_nnfunctions_flt.h:2038`
## arm_nn_mean_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mean_f32(
const float32_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Computes the mean of a float32 tensor along the specified axes.
Values are accumulated and divided once in float32; unlike the float16 variant there is no wider accumulator, so rounding error can grow with the reduction length, matching arm_reduce_sum_f32. Because the intermediate accumulation is itself float32, it can saturate to +/-Inf even when the mean itself is representable, but whether it does depends on accumulation order: a strictly sequential build keeps one running sum, while vector builds MVE intrinsics, or compiler auto-vectorization of the scalar path at -Ofast fold per-lane partial sums, so on inputs whose partial sums exceed FLT_MAX in magnitude either build may return +/-Inf and the two may disagree (one finite, one Inf); when partial sums of opposite sign both saturate, the vector fold can even yield NaN (Inf + -Inf) from all-finite inputs. Only when every accumulation order overflows e.g. same-signed values summing past FLT_MAX is +/-Inf guaranteed on all builds. This is the accumulation-order divergence described below taken to the extreme. NaN and Inf propagate. A mean over all -0.0f inputs returns +0.0f on every build: the accumulator starts at +0.0f and (+0.0f) + (-0.0f) is +0.0f under round-to-nearest. Vector and scalar builds may differ in final ulps because float accumulation order differs.
Unlike arm_reduce_sum_f32 (identical signature, null checks only), this kernel validates shapes and returns `ARM_CMSIS_NN_ARG_ERROR` when any input dimension is less than 1, when any `output_dims` entry differs from the input shape with the reduced axes collapsed to 1, or when the input element count or the reduction count does not fit in int32_t. `output_data` must not overlap `input_data:` each output element is written after reading its whole reduction set, so an aliased write can corrupt inputs still to be read.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (4D NHWC) |
| axis_dims | const cmsis_nn_dims * | in | 4D binary axis mask (non-zero = reduce that axis) |
| output_data | float32_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (reduced axes have size 1) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:2086`
## arm_reduce_sum_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_sum_f16(
const float16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Computes the sum of the input tensor along the specified axes.
Sums are accumulated in float32 (also for the float16 variant, which rounds once to float16 at the end), so results do not overflow at float16 range and precision does not degrade with the reduction count. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (4D NHWC) |
| axis_dims | const cmsis_nn_dims * | in | 4D binary axis mask (non-zero = reduce that axis) |
| output_data | float16_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (reduced axes have size 1) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3646`
## arm_argmin_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_argmin_f16(
const float16_t *input_data,
const cmsis_nn_dims *input_dims,
int32_t axis,
int32_t *output_data
)
```
Returns the first minimum's axis-relative INT32 index for a f16 tensor.
The input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.
A NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.
Metadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis | int32_t | in | Canonical reduction axis, in [0,3]. |
| output_data | int32_t * | out | INT32 indices, each in [0,input_dims[axis]). |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:3684`
## arm_argmax_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_argmax_f16(
const float16_t *input_data,
const cmsis_nn_dims *input_dims,
int32_t axis,
int32_t *output_data
)
```
Returns the first maximum's axis-relative INT32 index for a f16 tensor.
The input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.
A NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.
Metadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis | int32_t | in | Canonical reduction axis, in [0,3]. |
| output_data | int32_t * | out | INT32 indices, each in [0,input_dims[axis]). |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:3718`
## arm_reduce_max_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_max_f16(
const float16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Reduces a f16 NHWC tensor to its maximum along a binary axis mask.
Values are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7e00); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.
A zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces -Inf; an empty output performs no accesses to data buffers.
All metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis_dims | const cmsis_nn_dims * | in | Four binary reduction flags. |
| output_data | float16_t * | out | Output tensor. |
| output_dims | const cmsis_nn_dims * | in | NHWC output shape with reduced axes retained as 1. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write. |
Source: `Include/arm_nnfunctions_flt.h:3748`
## arm_reduce_min_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_min_f16(
const float16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Reduces a f16 NHWC tensor to its minimum along a binary axis mask.
Values are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7e00); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.
A zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces +Inf; an empty output performs no accesses to data buffers.
All metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis_dims | const cmsis_nn_dims * | in | Four binary reduction flags. |
| output_data | float16_t * | out | Output tensor. |
| output_dims | const cmsis_nn_dims * | in | NHWC output shape with reduced axes retained as 1. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write. |
Source: `Include/arm_nnfunctions_flt.h:3782`
## arm_nn_mean_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mean_f16(
const float16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Computes the mean of a float16 tensor along the specified axes.
Values are accumulated and divided in float32, then rounded once to float16. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs. Builds at -Ofast (the shipped CMSIS_OPTIMIZATION_LEVEL) may additionally differ from lower optimization levels by 1 ulp for non-power-of-two reduction counts: -freciprocal-math turns the divide-by-count into a multiply-by-reciprocal, which rounds differently.
Unlike arm_reduce_sum_f16 (identical signature, null checks only), this kernel validates shapes and returns `ARM_CMSIS_NN_ARG_ERROR` when any input dimension is less than 1, when any `output_dims` entry differs from the input shape with the reduced axes collapsed to 1, or when the input element count or the reduction count does not fit in int32_t. `output_data` must not overlap `input_data:` each output element is written after reading its whole reduction set, so an aliased write can corrupt inputs still to be read.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (4D NHWC) |
| axis_dims | const cmsis_nn_dims * | in | 4D binary axis mask (non-zero = reduce that axis) |
| output_data | float16_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (reduced axes have size 1) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3817`
---
# heliaCORE.Reshape
## arm_resize_nearest_neighbor_f32_get_buffer_size
`function` · `c`
```c
int32_t arm_resize_nearest_neighbor_f32_get_buffer_size(const cmsis_nn_dims *output_dims)
```
Scratch size in bytes for `arm_resize_nearest_neighbor_f32()` / `arm_resize_nearest_neighbor_f16()`.
The kernels precompute one int32_t input index per output row and per output column, so the requirement is (output_dims->h + output_dims->w) * sizeof(int32_t). Returns -1 (never 0) when `output_dims` is NULL, when h or w is less than 1, or when the size does not fit in int32_t; a negative result must not be used to size a buffer, and the kernels reject a { NULL, 0 } context outright (the -1 family of the integer sizers, not the 0-returning family most float sizers use; see the sentinel note on arm_nn_size_mul).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (only h and w are read). |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required ctx->size in bytes, or -1. |
Source: `Include/arm_nnfunctions_flt.h:1425`
## arm_resize_nearest_neighbor_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_resize_nearest_neighbor_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_resize_params *resize_params,
const cmsis_nn_dims *input_shape,
const float32_t *input_data,
const cmsis_nn_dims *output_size_shape,
const int32_t *output_size_data,
const cmsis_nn_dims *output_shape,
float32_t *output_data
)
```
Nearest-neighbor resize of a float32 NHWC tensor.
Pure data movement: every output element is a bit copy of one input element, so NaN (sign and payload), +/-Inf, -0.0 and subnormals are preserved bit-for-bit on the scalar and MVE legs alike (the MVE copy is a tail-predicated vldr/vstr pair, not an FP operation, so FPSCR flush-to-zero and default-NaN do not apply). No arithmetic is performed on the data; the only float math is the float32 index scale below.
Index semantics are TFLite's RESIZE_NEAREST_NEIGHBOR reference, evaluated in float32 per axis: scale = (align_corners && out > 1) ? (in - 1) / (out - 1) : in / out idx = align_corners ? roundf((o + offset) * scale) : floorf((o + offset) * scale) idx = min(idx, in - 1); if (half_pixel_centers) idx = max(idx, 0) with offset = half_pixel_centers ? 0.5f : 0.0f; roundf rounds ties away from zero, matching TfLiteRound. All four align_corners/half_pixel_centers combinations were verified bit-for-bit against TFLite 2.20 over a shape sweep (1..16 square, 80 random NHWC shapes up to 40x40, 224->7 and 7->224), including the out == 1 align_corners case, which maps to input index 0.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Scratch context. ctx->buf must be non-NULL and 4-byte aligned, and ctx->size at least arm_resize_nearest_neighbor_f32_get_buffer_size(output_shape); the kernel writes the x/y index maps here and does not read them after returning. |
| resize_params | const cmsis_nn_resize_params * | in | align_corners / half_pixel_centers. |
| input_shape | const cmsis_nn_dims * | in | Input tensor dimensions in NHWC format; every dimension must be >= 1. |
| input_data | const float32_t * | in | Input tensor data. Must not overlap `output_data`. |
| output_size_shape | const cmsis_nn_dims * | in | Dimensions of the output-size tensor; must hold exactly 2 elements. |
| output_size_data | const int32_t * | in | Output size as [output_height, output_width], both >= 1. |
| output_shape | const cmsis_nn_dims * | in | Output tensor dimensions in NHWC format; n and c must equal the input's and h/w must equal `output_size_data`. |
| output_data | float32_t * | out | Output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when any constraint above fails (including a NULL pointer argument); nothing is written on ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:1458`
## arm_resize_nearest_neighbor_f16_get_buffer_size
`function` · `c`
```c
int32_t arm_resize_nearest_neighbor_f16_get_buffer_size(const cmsis_nn_dims *output_dims)
```
Scratch size in bytes for `arm_resize_nearest_neighbor_f32()` / `arm_resize_nearest_neighbor_f16()`.
The kernels precompute one int32_t input index per output row and per output column, so the requirement is (output_dims->h + output_dims->w) * sizeof(int32_t). Returns -1 (never 0) when `output_dims` is NULL, when h or w is less than 1, or when the size does not fit in int32_t; a negative result must not be used to size a buffer, and the kernels reject a { NULL, 0 } context outright (the -1 family of the integer sizers, not the 0-returning family most float sizers use; see the sentinel note on arm_nn_size_mul).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (only h and w are read). |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required ctx->size in bytes, or -1. |
Source: `Include/arm_nnfunctions_flt.h:3294`
## arm_resize_nearest_neighbor_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_resize_nearest_neighbor_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_resize_params *resize_params,
const cmsis_nn_dims *input_shape,
const float16_t *input_data,
const cmsis_nn_dims *output_size_shape,
const int32_t *output_size_data,
const cmsis_nn_dims *output_shape,
float16_t *output_data
)
```
Nearest-neighbor resize of a float32 NHWC tensor.
Pure data movement: every output element is a bit copy of one input element, so NaN (sign and payload), +/-Inf, -0.0 and subnormals are preserved bit-for-bit on the scalar and MVE legs alike (the MVE copy is a tail-predicated vldr/vstr pair, not an FP operation, so FPSCR flush-to-zero and default-NaN do not apply). No arithmetic is performed on the data; the only float math is the float32 index scale below.
Index semantics are TFLite's RESIZE_NEAREST_NEIGHBOR reference, evaluated in float32 per axis: scale = (align_corners && out > 1) ? (in - 1) / (out - 1) : in / out idx = align_corners ? roundf((o + offset) * scale) : floorf((o + offset) * scale) idx = min(idx, in - 1); if (half_pixel_centers) idx = max(idx, 0) with offset = half_pixel_centers ? 0.5f : 0.0f; roundf rounds ties away from zero, matching TfLiteRound. All four align_corners/half_pixel_centers combinations were verified bit-for-bit against TFLite 2.20 over a shape sweep (1..16 square, 80 random NHWC shapes up to 40x40, 224->7 and 7->224), including the out == 1 align_corners case, which maps to input index 0.
:::note
float16 twin: each element is copied as a 16-bit lane with no widening or conversion, so half-precision NaN payloads and subnormals are preserved exactly and the data is never evaluated in float32. Scratch is sized by `arm_resize_nearest_neighbor_f16_get_buffer_size()` (same query as f32).
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Scratch context. ctx->buf must be non-NULL and 4-byte aligned, and ctx->size at least arm_resize_nearest_neighbor_f32_get_buffer_size(output_shape); the kernel writes the x/y index maps here and does not read them after returning. |
| resize_params | const cmsis_nn_resize_params * | in | align_corners / half_pixel_centers. |
| input_shape | const cmsis_nn_dims * | in | Input tensor dimensions in NHWC format; every dimension must be >= 1. |
| input_data | const float16_t * | in | Input tensor data. Must not overlap `output_data`. |
| output_size_shape | const cmsis_nn_dims * | in | Dimensions of the output-size tensor; must hold exactly 2 elements. |
| output_size_data | const int32_t * | in | Output size as [output_height, output_width], both >= 1. |
| output_shape | const cmsis_nn_dims * | in | Output tensor dimensions in NHWC format; n and c must equal the input's and h/w must equal `output_size_data`. |
| output_data | float16_t * | out | Output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when any constraint above fails (including a NULL pointer argument); nothing is written on ARG_ERROR. |
Source: `Include/arm_nnfunctions_flt.h:3302`
---
# heliaCORE.Softmax
## arm_softmax_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_softmax_f32(const float32_t *input, int32_t num_rows, int32_t row_size, float32_t *output)
```
Softmax using the float-native API signature.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input matrix stored as `num_rows` rows of `row_size` values. |
| num_rows | int32_t | in | Number of rows in the input matrix. |
| row_size | int32_t | in | Number of columns per row. |
| output | float32_t * | out | Pointer to the output matrix. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1882`
## arm_softmax_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_softmax_f16(const float16_t *input, int32_t num_rows, int32_t row_size, float16_t *output)
```
Softmax using the float-native API signature.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input matrix stored as `num_rows` rows of `row_size` values. |
| num_rows | int32_t | in | Number of rows in the input matrix. |
| row_size | int32_t | in | Number of columns per row. |
| output | float16_t * | out | Pointer to the output matrix. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3640`
---
# heliaCORE.StridedSlice
## arm_strided_slice_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_strided_slice_f32(
const float32_t *input_data,
float32_t *output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const begin_dims,
const cmsis_nn_dims *const stride_dims,
const cmsis_nn_dims *const output_dims
)
```
Strided slice for float32 data (pure copy, TensorFlow Lite compatible).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Pointer to input tensor. |
| output_data | float32_t * | out | Pointer to output tensor. |
| input_dims | const cmsis_nn_dims *const | in | Input tensor dimensions. |
| begin_dims | const cmsis_nn_dims *const | in | Begin dimensions for slicing. |
| stride_dims | const cmsis_nn_dims *const | in | Stride dimensions for slicing. |
| output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success. |
Source: `Include/arm_nnfunctions_flt.h:823`
---
# heliaCORE.supportConversion
// end group groupPrivTypes
Perform data type conversion in-between neural network operations
---
# heliaCORE.SVDF
## arm_svdf_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_svdf_f32(
const cmsis_nn_context *ctx,
const cmsis_nn_context *input_ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_svdf_params_f32 *svdf_params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *state_dims,
float32_t *state_data,
const cmsis_nn_dims *weights_feature_dims,
const float32_t *weights_feature_data,
const cmsis_nn_dims *weights_time_dims,
const float32_t *weights_time_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
Stateful singular value decomposition filter.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused by this function. Reserved for future use; may be NULL. |
| input_ctx | const cmsis_nn_context * | in, out | Mandatory, not optional: a NULL input_ctx, or a NULL input_ctx->buf, is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Staging buffer written by this function, holding one element per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Sized by arm_svdf_f32_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(float32_t) bytes. Setting input_ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check. The caller is expected to clear the buffer, if applicable, for security reasons. |
| output_ctx | const cmsis_nn_context * | in, out | Mandatory, not optional: a NULL output_ctx, or a NULL output_ctx->buf, is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Staging buffer written by this function, holding one element per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Sized by arm_svdf_f32_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(float32_t) bytes, truncating division. Setting output_ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check. The caller is expected to clear the buffer, if applicable, for security reasons. |
| svdf_params | const cmsis_nn_svdf_params_f32 * | in | SVDF operator parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| state_dims | const cmsis_nn_dims * | in | State tensor dimensions. |
| state_data | float32_t * | in, out | Pointer to the mutable state tensor. |
| weights_feature_dims | const cmsis_nn_dims * | in | Feature-weight tensor dimensions. |
| weights_feature_data | const float32_t * | in | Pointer to the feature-weight tensor. |
| weights_time_dims | const cmsis_nn_dims * | in | Time-weight tensor dimensions. |
| weights_time_data | const float32_t * | in | Pointer to the time-weight tensor. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:1694`
## arm_svdf_f32_input_ctx_get_buffer_size
`function` · `c`
```c
int32_t arm_svdf_f32_input_ctx_get_buffer_size(
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)
```
Get size of the input_ctx staging buffer required by `arm_svdf_f32()`.
:::note
This query reports an out-of-range shape as -1, following the SVDF family (`arm_svdf_s8_get_buffer_size()`). That differs from the float convolution and fully-connected queries in this header, which report an out-of-range size as 0. The reason is that `arm_svdf_f32()` reads ctx->size, and size == 0 is the opt-out signal for its scratch-size check: a 0-on-overflow answer fed straight back as `buf = alloc(0), size = 0` would silently disable the check over a zero-byte allocation, whereas alloc((size_t)-1) fails and the NULL check catches it.
:::
:::note
0 is still a valid *return* for a degenerate shape (input_dims->n == 0). Unlike the general rule in README.md, a 0 here does NOT mean you may pass { NULL, 0 }: `arm_svdf_f32()` rejects a NULL input_ctx->buf with ARM_CMSIS_NN_ARG_ERROR regardless of the size. Allocate a non-NULL pointer, or do not call the kernel for a shape that produces no output.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f32()`. |
| weights_feature_dims | const cmsis_nn_dims * | in | Feature-weight tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f32()`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: input_dims->n * weights_feature_dims->n * sizeof(float32_t). Returns -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the product would not fit in an int32_t. The figure and the validation are the same on every build target, since `arm_svdf_f32()` stages this buffer on every build rather than only under MVE. |
Source: `Include/arm_nnfunctions_flt.h:1734`
## arm_svdf_f32_output_ctx_get_buffer_size
`function` · `c`
```c
int32_t arm_svdf_f32_output_ctx_get_buffer_size(
const cmsis_nn_svdf_params_f32 *svdf_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)
```
Get size of the output_ctx staging buffer required by `arm_svdf_f32()`.
:::note
Same -1 and degenerate-0 contract as `arm_svdf_f32_input_ctx_get_buffer_size()`, including that a 0 does not license passing { NULL, 0 }.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| svdf_params | const cmsis_nn_svdf_params_f32 * | in | SVDF operator parameters; only svdf_params->rank is read. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f32()`. |
| weights_feature_dims | const cmsis_nn_dims * | in | Feature-weight tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f32()`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(float32_t), truncating division to match the kernel's own unit count. Returns -1 if any pointer is NULL, if svdf_params->rank is zero or negative, if input_dims->n or weights_feature_dims->n is negative, or if the product would not fit in an int32_t. |
Source: `Include/arm_nnfunctions_flt.h:1754`
## arm_svdf_f16
`function` · `c`
```c
arm_cmsis_nn_status arm_svdf_f16(
const cmsis_nn_context *ctx,
const cmsis_nn_context *input_ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_svdf_params_f16 *svdf_params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *state_dims,
float16_t *state_data,
const cmsis_nn_dims *weights_feature_dims,
const float16_t *weights_feature_data,
const cmsis_nn_dims *weights_time_dims,
const float16_t *weights_time_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
Stateful singular value decomposition filter, float16 variant.
:::note
Sizing an f16 layer with the `arm_svdf_f32()` queries over-allocates and is safe. Sizing an f32 layer with the `arm_svdf_f16()` queries under-allocates by half: `arm_svdf_f32()` returns ARM_CMSIS_NN_ARG_ERROR if ctx->size carries that undersized figure, but corrupts memory if ctx->size is left at 0, which opts out of the check.
:::
:::note
NaN propagates through the activation clamps that take the bit-classified scalar clamp of #380, at every optimization level on the gated toolchains, including the shipped -Ofast: the input-activation clamp is that scalar clamp on EVERY build, and the output-activation clamp is on non-MVE builds. On MVE builds the output-activation clamp is vmaxnmq/vminnmq with no NaN restore, so a NaN resolves to a clamp bound there instead.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused by this function. Reserved for future use; may be NULL. |
| input_ctx | const cmsis_nn_context * | in, out | Mandatory, not optional: a NULL input_ctx, or a NULL input_ctx->buf, is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Staging buffer written by this function, holding one element per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Sized by arm_svdf_f16_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(float16_t) bytes. Note this is float16_t, half the `arm_svdf_f32()` figure for the same shape. Setting input_ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check. The caller is expected to clear the buffer, if applicable, for security reasons. |
| output_ctx | const cmsis_nn_context * | in, out | Mandatory, not optional: a NULL output_ctx, or a NULL output_ctx->buf, is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Staging buffer written by this function, holding one element per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Sized by arm_svdf_f16_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(float16_t) bytes, truncating division. Note this is float16_t, half the `arm_svdf_f32()` figure for the same shape. Setting output_ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check. The caller is expected to clear the buffer, if applicable, for security reasons. |
| svdf_params | const cmsis_nn_svdf_params_f16 * | in | SVDF operator parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| state_dims | const cmsis_nn_dims * | in | State tensor dimensions. |
| state_data | float16_t * | in, out | Pointer to the mutable state tensor. |
| weights_feature_dims | const cmsis_nn_dims * | in | Feature-weight tensor dimensions. |
| weights_feature_data | const float16_t * | in | Pointer to the feature-weight tensor. |
| weights_time_dims | const cmsis_nn_dims * | in | Time-weight tensor dimensions. |
| weights_time_data | const float16_t * | in | Pointer to the time-weight tensor. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |
Source: `Include/arm_nnfunctions_flt.h:3479`
## arm_svdf_f16_input_ctx_get_buffer_size
`function` · `c`
```c
int32_t arm_svdf_f16_input_ctx_get_buffer_size(
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)
```
Get size of the input_ctx staging buffer required by `arm_svdf_f16()`.
:::note
`arm_svdf_f16()` stages float16_t, so this is HALF the byte count `arm_svdf_f32_input_ctx_get_buffer_size()` returns for the same shape. Sizing an f16 layer with the f32 query over-allocates and is safe; sizing an f32 layer with this one under-allocates by half.
:::
:::note
This query reports an out-of-range shape as -1, following the SVDF family (`arm_svdf_s8_get_buffer_size()`), not the 0 used by the float convolution and fully-connected queries in this header. The reason is that `arm_svdf_f16()` reads ctx->size, and size == 0 is the opt-out signal for its scratch-size check: a 0-on-overflow answer fed straight back as `buf = alloc(0), size = 0` would silently disable the check over a zero-byte allocation, whereas alloc((size_t)-1) fails and the NULL check catches it.
:::
:::note
0 is still a valid *return* for a degenerate shape (input_dims->n == 0). Unlike the general rule in README.md, a 0 here does NOT mean you may pass { NULL, 0 }: `arm_svdf_f16()` rejects a NULL input_ctx->buf with ARM_CMSIS_NN_ARG_ERROR regardless of the size.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f16()`. |
| weights_feature_dims | const cmsis_nn_dims * | in | Feature-weight tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f16()`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: input_dims->n * weights_feature_dims->n * sizeof(float16_t). Returns -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the product would not fit in an int32_t. The figure and the validation are the same on every build target, since `arm_svdf_f16()` stages this buffer on every build rather than only under MVE. |
Source: `Include/arm_nnfunctions_flt.h:3521`
## arm_svdf_f16_output_ctx_get_buffer_size
`function` · `c`
```c
int32_t arm_svdf_f16_output_ctx_get_buffer_size(
const cmsis_nn_svdf_params_f16 *svdf_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)
```
Get size of the output_ctx staging buffer required by `arm_svdf_f16()`.
:::note
`arm_svdf_f16()` stages float16_t, so this is HALF the byte count `arm_svdf_f32_output_ctx_get_buffer_size()` returns for the same shape.
:::
:::note
Same -1 and degenerate-0 contract as `arm_svdf_f16_input_ctx_get_buffer_size()`, including that a 0 does not license passing { NULL, 0 }. A rank greater than weights_feature_dims->n truncates the unit count to 0 and so returns 0.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| svdf_params | const cmsis_nn_svdf_params_f16 * | in | SVDF operator parameters; only svdf_params->rank is read. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f16()`. |
| weights_feature_dims | const cmsis_nn_dims * | in | Feature-weight tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f16()`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(float16_t), truncating division to match the kernel's own unit count. Returns -1 if any pointer is NULL, if svdf_params->rank is zero or negative, if input_dims->n or weights_feature_dims->n is negative, or if the product would not fit in an int32_t. |
Source: `Include/arm_nnfunctions_flt.h:3544`
---
# heliaCORE.Internal
---
# heliaCORE.Internal.arm_concatenation_common
## ARM_CONCATENATION_DEFINE
`macro` · `c`
```c
#define ARM_CONCATENATION_DEFINE(SUFFIX, TYPE)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| SUFFIX | | | |
| TYPE | | | |
Source: `Include/Internal/arm_concatenation_common.h:36`
---
# heliaCORE.Internal.arm_conv_opt_common
## ARM_CONV_SPEC_ENTRY
`macro` · `c`
```c
#define ARM_CONV_SPEC_ENTRY(MATCH_FN, CALL_FN) {(MATCH_FN), (CALL_FN)}
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| MATCH_FN | | | |
| CALL_FN | | | |
Source: `Include/Internal/arm_conv_opt_common.h:34`
## ARM_CONV_ARRAY_SIZE
`macro` · `c`
```c
#define ARM_CONV_ARRAY_SIZE(arr) (sizeof(arr) / sizeof((arr)[0]))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| arr | | | |
Source: `Include/Internal/arm_conv_opt_common.h:36`
## ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MAX_TILE_ROWS
`macro` · `c`
```c
#define ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MAX_TILE_ROWS (8)
```
Source: `Include/Internal/arm_conv_opt_common.h:45`
## ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_K
`macro` · `c`
```c
#define ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_K (16)
```
Source: `Include/Internal/arm_conv_opt_common.h:46`
## ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_OC
`macro` · `c`
```c
#define ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_OC (8)
```
Source: `Include/Internal/arm_conv_opt_common.h:47`
## ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_POS
`macro` · `c`
```c
#define ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_POS (8)
```
Source: `Include/Internal/arm_conv_opt_common.h:48`
## ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MAX_TILE_ROWS
`macro` · `c`
```c
#define ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MAX_TILE_ROWS (8)
```
Source: `Include/Internal/arm_conv_opt_common.h:54`
## ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_K
`macro` · `c`
```c
#define ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_K (16)
```
Source: `Include/Internal/arm_conv_opt_common.h:55`
## ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_OC
`macro` · `c`
```c
#define ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_OC (8)
```
Source: `Include/Internal/arm_conv_opt_common.h:56`
## ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_POS
`macro` · `c`
```c
#define ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_POS (8)
```
Source: `Include/Internal/arm_conv_opt_common.h:57`
## ARM_CONV_DISPATCH
`macro` · `c`
```c
#define ARM_CONV_DISPATCH(TABLE, COUNT, ...) do \ { \ for (size_t _i = 0; _i < (COUNT); ++_i) \ { \ if ((TABLE)[_i].match(__VA_ARGS__)) \ { \ return (TABLE)[_i].call(__VA_ARGS__); \ } \ } \ } while (0)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| TABLE | | | |
| COUNT | | | |
| ... | | | |
Source: `Include/Internal/arm_conv_opt_common.h:59`
---
# heliaCORE.Internal.arm_conv_opt_f16
## arm_conv_spec_f16
`struct` · `c`
```c
struct arm_conv_spec_f16
```
Source: `Include/Internal/arm_conv_opt_f16.h:63`
### arm_conv_spec_f16::match
`attribute` · `c`
```c
arm_conv_match_f16 match
```
Source: `Include/Internal/arm_conv_opt_f16.h:65`
### arm_conv_spec_f16::call
`attribute` · `c`
```c
arm_conv_call_f16 call
```
Source: `Include/Internal/arm_conv_opt_f16.h:66`
## arm_conv_match_f16
`type` · `c`
```c
typedef bool(* arm_conv_match_f16) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input_data, const cmsis_nn_dims *filter_dims, const float16_t *filter_data, const cmsis_nn_dims *bias_dims, const float16_t *bias_data, const cmsis_nn_dims *output_dims, float16_t *output_data)
```
Source: `Include/Internal/arm_conv_opt_f16.h:41`
## arm_conv_call_f16
`type` · `c`
```c
typedef arm_cmsis_nn_status(* arm_conv_call_f16) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input_data, const cmsis_nn_dims *filter_dims, const float16_t *filter_data, const cmsis_nn_dims *bias_dims, const float16_t *bias_data, const cmsis_nn_dims *output_dims, float16_t *output_data)
```
Source: `Include/Internal/arm_conv_opt_f16.h:52`
## arm_conv1d_spec_k5_nhwc_f16_match
`function` · `c`
```c
static bool arm_conv1d_spec_k5_nhwc_f16_match(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float16_t * | | |
Source: `Include/Internal/arm_conv_opt_f16.h:69`
## arm_conv1d_spec_k5_nhwc_f16_call_body
`function` · `c`
```c
static arm_cmsis_nn_status arm_conv1d_spec_k5_nhwc_f16_call_body(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
const bool acc16
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float16_t * | | |
| acc16 | const bool | | |
Source: `Include/Internal/arm_conv_opt_f16.h:103`
## arm_conv1d_spec_k5_nhwc_f16_call
`function` · `c`
```c
static arm_cmsis_nn_status arm_conv1d_spec_k5_nhwc_f16_call(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float16_t * | | |
Source: `Include/Internal/arm_conv_opt_f16.h:141`
## arm_conv1d_spec_k5_nhwc_f16_call_acc16
`function` · `c`
```c
static arm_cmsis_nn_status arm_conv1d_spec_k5_nhwc_f16_call_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float16_t * | | |
Source: `Include/Internal/arm_conv_opt_f16.h:165`
## arm_conv1d_spec_k3_nhwc_f16_match
`function` · `c`
```c
static bool arm_conv1d_spec_k3_nhwc_f16_match(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float16_t * | | |
Source: `Include/Internal/arm_conv_opt_f16.h:189`
## arm_conv1d_spec_k3_nhwc_f16_call_body
`function` · `c`
```c
static arm_cmsis_nn_status arm_conv1d_spec_k3_nhwc_f16_call_body(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data,
const bool acc16
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float16_t * | | |
| acc16 | const bool | | |
Source: `Include/Internal/arm_conv_opt_f16.h:223`
## arm_conv1d_spec_k3_nhwc_f16_call
`function` · `c`
```c
static arm_cmsis_nn_status arm_conv1d_spec_k3_nhwc_f16_call(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float16_t * | | |
Source: `Include/Internal/arm_conv_opt_f16.h:261`
## arm_conv1d_spec_k3_nhwc_f16_call_acc16
`function` · `c`
```c
static arm_cmsis_nn_status arm_conv1d_spec_k3_nhwc_f16_call_acc16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input_data,
const cmsis_nn_dims *filter_dims,
const float16_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float16_t *bias_data,
const cmsis_nn_dims *output_dims,
float16_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float16_t * | | |
Source: `Include/Internal/arm_conv_opt_f16.h:285`
## arm_conv_spec_nhwc_f16
`attribute` · `c`
```c
const arm_conv_spec_f16 arm_conv_spec_nhwc_f16[]
```
Source: `Include/Internal/arm_conv_opt_f16.h:310`
## arm_conv_spec_nhwc_f16_acc16
`attribute` · `c`
```c
const arm_conv_spec_f16 arm_conv_spec_nhwc_f16_acc16[]
```
Source: `Include/Internal/arm_conv_opt_f16.h:317`
## arm_conv_spec_nhwc_f16_matches_any
`function` · `c`
```c
static bool arm_conv_spec_nhwc_f16_matches_any(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| filter_dims | const cmsis_nn_dims * | | |
| output_dims | const cmsis_nn_dims * | | |
Source: `Include/Internal/arm_conv_opt_f16.h:326`
---
# heliaCORE.Internal.arm_conv_opt_f32
## arm_conv_spec_f32
`struct` · `c`
```c
struct arm_conv_spec_f32
```
Source: `Include/Internal/arm_conv_opt_f32.h:63`
### arm_conv_spec_f32::match
`attribute` · `c`
```c
arm_conv_match_f32 match
```
Source: `Include/Internal/arm_conv_opt_f32.h:65`
### arm_conv_spec_f32::call
`attribute` · `c`
```c
arm_conv_call_f32 call
```
Source: `Include/Internal/arm_conv_opt_f32.h:66`
## arm_conv_match_f32
`type` · `c`
```c
typedef bool(* arm_conv_match_f32) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input_data, const cmsis_nn_dims *filter_dims, const float32_t *filter_data, const cmsis_nn_dims *bias_dims, const float32_t *bias_data, const cmsis_nn_dims *output_dims, float32_t *output_data)
```
Source: `Include/Internal/arm_conv_opt_f32.h:41`
## arm_conv_call_f32
`type` · `c`
```c
typedef arm_cmsis_nn_status(* arm_conv_call_f32) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input_data, const cmsis_nn_dims *filter_dims, const float32_t *filter_data, const cmsis_nn_dims *bias_dims, const float32_t *bias_data, const cmsis_nn_dims *output_dims, float32_t *output_data)
```
Source: `Include/Internal/arm_conv_opt_f32.h:52`
## arm_conv1d_spec_k5_nhwc_f32_match
`function` · `c`
```c
static bool arm_conv1d_spec_k5_nhwc_f32_match(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f32 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float32_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float32_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float32_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float32_t * | | |
Source: `Include/Internal/arm_conv_opt_f32.h:69`
## arm_conv1d_spec_k5_nhwc_f32_call
`function` · `c`
```c
static arm_cmsis_nn_status arm_conv1d_spec_k5_nhwc_f32_call(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f32 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float32_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float32_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float32_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float32_t * | | |
Source: `Include/Internal/arm_conv_opt_f32.h:103`
## arm_conv1d_spec_k3_nhwc_f32_match
`function` · `c`
```c
static bool arm_conv1d_spec_k3_nhwc_f32_match(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f32 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float32_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float32_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float32_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float32_t * | | |
Source: `Include/Internal/arm_conv_opt_f32.h:140`
## arm_conv1d_spec_k3_nhwc_f32_call
`function` · `c`
```c
static arm_cmsis_nn_status arm_conv1d_spec_k3_nhwc_f32_call(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *params,
const cmsis_nn_dims *input_dims,
const float32_t *input_data,
const cmsis_nn_dims *filter_dims,
const float32_t *filter_data,
const cmsis_nn_dims *bias_dims,
const float32_t *bias_data,
const cmsis_nn_dims *output_dims,
float32_t *output_data
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f32 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input_data | const float32_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| filter_data | const float32_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias_data | const float32_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output_data | float32_t * | | |
Source: `Include/Internal/arm_conv_opt_f32.h:174`
## arm_conv_spec_nhwc_f32
`attribute` · `c`
```c
const arm_conv_spec_f32 arm_conv_spec_nhwc_f32[]
```
Source: `Include/Internal/arm_conv_opt_f32.h:211`
## arm_conv_spec_nhwc_f32_matches_any
`function` · `c`
```c
static bool arm_conv_spec_nhwc_f32_matches_any(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params_f32 *params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_conv_params_f32 * | | |
| input_dims | const cmsis_nn_dims * | | |
| filter_dims | const cmsis_nn_dims * | | |
| output_dims | const cmsis_nn_dims * | | |
Source: `Include/Internal/arm_conv_opt_f32.h:216`
---
# heliaCORE.Internal.arm_conv1x1_opt_common
## ARM_CONV1X1_SPEC_ENTRY
`macro` · `c`
```c
#define ARM_CONV1X1_SPEC_ENTRY(MATCH_FN, CALL_FN) { \ (MATCH_FN), (CALL_FN) \ }
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| MATCH_FN | | | |
| CALL_FN | | | |
Source: `Include/Internal/arm_conv1x1_opt_common.h:34`
## ARM_CONV1X1_ARRAY_SIZE
`macro` · `c`
```c
#define ARM_CONV1X1_ARRAY_SIZE(arr) (sizeof(arr) / sizeof((arr)[0]))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| arr | | | |
Source: `Include/Internal/arm_conv1x1_opt_common.h:39`
## ARM_CONV1X1_DISPATCH
`macro` · `c`
```c
#define ARM_CONV1X1_DISPATCH(TABLE, COUNT, ...) do \ { \ for (size_t _i = 0; _i < (COUNT); ++_i) \ { \ if ((TABLE)[_i].match(__VA_ARGS__)) \ { \ return (TABLE)[_i].call(__VA_ARGS__); \ } \ } \ } while (0)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| TABLE | | | |
| COUNT | | | |
| ... | | | |
Source: `Include/Internal/arm_conv1x1_opt_common.h:41`
---
# heliaCORE.Internal.arm_conv1x1_opt_f16
## arm_conv1x1_spec_f16
`struct` · `c`
```c
struct arm_conv1x1_spec_f16
```
Source: `Include/Internal/arm_conv1x1_opt_f16.h:62`
### arm_conv1x1_spec_f16::match
`attribute` · `c`
```c
arm_conv1x1_match_f16 match
```
Source: `Include/Internal/arm_conv1x1_opt_f16.h:64`
### arm_conv1x1_spec_f16::call
`attribute` · `c`
```c
arm_conv1x1_call_f16 call
```
Source: `Include/Internal/arm_conv1x1_opt_f16.h:65`
## arm_conv1x1_match_f16
`type` · `c`
```c
typedef bool(* arm_conv1x1_match_f16) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input_data, const cmsis_nn_dims *filter_dims, const float16_t *filter_data, const cmsis_nn_dims *bias_dims, const float16_t *bias_data, const cmsis_nn_dims *output_dims, float16_t *output_data)
```
Source: `Include/Internal/arm_conv1x1_opt_f16.h:40`
## arm_conv1x1_call_f16
`type` · `c`
```c
typedef arm_cmsis_nn_status(* arm_conv1x1_call_f16) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input_data, const cmsis_nn_dims *filter_dims, const float16_t *filter_data, const cmsis_nn_dims *bias_dims, const float16_t *bias_data, const cmsis_nn_dims *output_dims, float16_t *output_data)
```
Source: `Include/Internal/arm_conv1x1_opt_f16.h:51`
---
# heliaCORE.Internal.arm_conv1x1_opt_f32
## arm_conv1x1_spec_f32
`struct` · `c`
```c
struct arm_conv1x1_spec_f32
```
Source: `Include/Internal/arm_conv1x1_opt_f32.h:62`
### arm_conv1x1_spec_f32::match
`attribute` · `c`
```c
arm_conv1x1_match_f32 match
```
Source: `Include/Internal/arm_conv1x1_opt_f32.h:64`
### arm_conv1x1_spec_f32::call
`attribute` · `c`
```c
arm_conv1x1_call_f32 call
```
Source: `Include/Internal/arm_conv1x1_opt_f32.h:65`
## arm_conv1x1_match_f32
`type` · `c`
```c
typedef bool(* arm_conv1x1_match_f32) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input_data, const cmsis_nn_dims *filter_dims, const float32_t *filter_data, const cmsis_nn_dims *bias_dims, const float32_t *bias_data, const cmsis_nn_dims *output_dims, float32_t *output_data)
```
Source: `Include/Internal/arm_conv1x1_opt_f32.h:40`
## arm_conv1x1_call_f32
`type` · `c`
```c
typedef arm_cmsis_nn_status(* arm_conv1x1_call_f32) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input_data, const cmsis_nn_dims *filter_dims, const float32_t *filter_data, const cmsis_nn_dims *bias_dims, const float32_t *bias_data, const cmsis_nn_dims *output_dims, float32_t *output_data)
```
Source: `Include/Internal/arm_conv1x1_opt_f32.h:51`
---
# heliaCORE.Internal.arm_depthwise_conv_opt_common
## ARM_DW_SPEC_ENTRY
`macro` · `c`
```c
#define ARM_DW_SPEC_ENTRY(MATCH_FN, CALL_FN) { \ (MATCH_FN), (CALL_FN) \ }
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| MATCH_FN | | | |
| CALL_FN | | | |
Source: `Include/Internal/arm_depthwise_conv_opt_common.h:36`
## ARM_DW_ARRAY_SIZE
`macro` · `c`
```c
#define ARM_DW_ARRAY_SIZE(arr) (sizeof(arr) / sizeof((arr)[0]))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| arr | | | |
Source: `Include/Internal/arm_depthwise_conv_opt_common.h:41`
## arm_depthwise_conv_input_index_nhwc
`function` · `c`
```c
static int32_t arm_depthwise_conv_input_index_nhwc(int32_t x, int32_t y, int32_t c, int32_t input_x, int32_t input_ch)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | int32_t | | |
| y | int32_t | | |
| c | int32_t | | |
| input_x | int32_t | | |
| input_ch | int32_t | | |
Source: `Include/Internal/arm_depthwise_conv_opt_common.h:45`
## arm_depthwise_conv_output_index_nhwc
`function` · `c`
```c
static int32_t arm_depthwise_conv_output_index_nhwc(
int32_t out_x,
int32_t out_y,
int32_t out_ch,
int32_t output_x,
int32_t output_ch
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| out_x | int32_t | | |
| out_y | int32_t | | |
| out_ch | int32_t | | |
| output_x | int32_t | | |
| output_ch | int32_t | | |
Source: `Include/Internal/arm_depthwise_conv_opt_common.h:52`
## ARM_DW_DISPATCH
`macro` · `c`
```c
#define ARM_DW_DISPATCH(TABLE, COUNT, ...) do \ { \ for (size_t _i = 0; _i < (COUNT); ++_i) \ { \ if ((TABLE)[_i].match(__VA_ARGS__)) \ { \ return (TABLE)[_i].call(__VA_ARGS__); \ } \ } \ } while (0)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| TABLE | | | |
| COUNT | | | |
| ... | | | |
Source: `Include/Internal/arm_depthwise_conv_opt_common.h:57`
---
# heliaCORE.Internal.arm_depthwise_conv_opt_f16
## arm_dw_spec_f16
`struct` · `c`
```c
struct arm_dw_spec_f16
```
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:67`
### arm_dw_spec_f16::match
`attribute` · `c`
```c
arm_dw_match_f16 match
```
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:69`
### arm_dw_spec_f16::call
`attribute` · `c`
```c
arm_dw_call_f16 call
```
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:70`
## arm_dw_match_f16
`type` · `c`
```c
typedef bool(* arm_dw_match_f16) (const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input, const cmsis_nn_dims *filter_dims, const float16_t *kernel, const cmsis_nn_dims *bias_dims, const float16_t *bias, const cmsis_nn_dims *output_dims, float16_t *output, arm_nn_dw_kernel_layout_f16 kernel_layout)
```
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:43`
## arm_dw_call_f16
`type` · `c`
```c
typedef arm_cmsis_nn_status(* arm_dw_call_f16) (const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input, const cmsis_nn_dims *filter_dims, const float16_t *kernel, const cmsis_nn_dims *bias_dims, const float16_t *bias, const cmsis_nn_dims *output_dims, float16_t *output, arm_nn_dw_kernel_layout_f16 kernel_layout)
```
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:55`
## arm_dw_spec_k3_1d_nhwc_f16_match
`function` · `c`
```c
static bool arm_dw_spec_k3_1d_nhwc_f16_match(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output,
arm_nn_dw_kernel_layout_f16 kernel_layout
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_dw_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| kernel | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output | float16_t * | | |
| kernel_layout | arm_nn_dw_kernel_layout_f16 | | |
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:73`
## arm_dw_spec_k3_1d_nhwc_f16_call
`function` · `c`
```c
static arm_cmsis_nn_status arm_dw_spec_k3_1d_nhwc_f16_call(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output,
arm_nn_dw_kernel_layout_f16 kernel_layout
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_dw_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| kernel | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output | float16_t * | | |
| kernel_layout | arm_nn_dw_kernel_layout_f16 | | |
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:105`
## arm_dw_spec_2x5_nhwc_f16_match
`function` · `c`
```c
static bool arm_dw_spec_2x5_nhwc_f16_match(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output,
arm_nn_dw_kernel_layout_f16 kernel_layout
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_dw_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| kernel | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output | float16_t * | | |
| kernel_layout | arm_nn_dw_kernel_layout_f16 | | |
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:134`
## arm_dw_spec_2x5_nhwc_f16_call
`function` · `c`
```c
static arm_cmsis_nn_status arm_dw_spec_2x5_nhwc_f16_call(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f16 *params,
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *filter_dims,
const float16_t *kernel,
const cmsis_nn_dims *bias_dims,
const float16_t *bias,
const cmsis_nn_dims *output_dims,
float16_t *output,
arm_nn_dw_kernel_layout_f16 kernel_layout
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_dw_conv_params_f16 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input | const float16_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| kernel | const float16_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias | const float16_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output | float16_t * | | |
| kernel_layout | arm_nn_dw_kernel_layout_f16 | | |
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:163`
## arm_dw_spec_nhwc_f16
`attribute` · `c`
```c
const arm_dw_spec_f16 arm_dw_spec_nhwc_f16[]
```
Source: `Include/Internal/arm_depthwise_conv_opt_f16.h:199`
---
# heliaCORE.Internal.arm_depthwise_conv_opt_f32
## arm_dw_spec_f32
`struct` · `c`
```c
struct arm_dw_spec_f32
```
Source: `Include/Internal/arm_depthwise_conv_opt_f32.h:67`
### arm_dw_spec_f32::match
`attribute` · `c`
```c
arm_dw_match_f32 match
```
Source: `Include/Internal/arm_depthwise_conv_opt_f32.h:69`
### arm_dw_spec_f32::call
`attribute` · `c`
```c
arm_dw_call_f32 call
```
Source: `Include/Internal/arm_depthwise_conv_opt_f32.h:70`
## arm_dw_match_f32
`type` · `c`
```c
typedef bool(* arm_dw_match_f32) (const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input, const cmsis_nn_dims *filter_dims, const float32_t *kernel, const cmsis_nn_dims *bias_dims, const float32_t *bias, const cmsis_nn_dims *output_dims, float32_t *output, arm_nn_dw_kernel_layout_f32 kernel_layout)
```
Source: `Include/Internal/arm_depthwise_conv_opt_f32.h:43`
## arm_dw_call_f32
`type` · `c`
```c
typedef arm_cmsis_nn_status(* arm_dw_call_f32) (const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input, const cmsis_nn_dims *filter_dims, const float32_t *kernel, const cmsis_nn_dims *bias_dims, const float32_t *bias, const cmsis_nn_dims *output_dims, float32_t *output, arm_nn_dw_kernel_layout_f32 kernel_layout)
```
Source: `Include/Internal/arm_depthwise_conv_opt_f32.h:55`
## arm_dw_spec_k3_1d_nhwc_f32_match
`function` · `c`
```c
static bool arm_dw_spec_k3_1d_nhwc_f32_match(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f32 *params,
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *filter_dims,
const float32_t *kernel,
const cmsis_nn_dims *bias_dims,
const float32_t *bias,
const cmsis_nn_dims *output_dims,
float32_t *output,
arm_nn_dw_kernel_layout_f32 kernel_layout
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_dw_conv_params_f32 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input | const float32_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| kernel | const float32_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias | const float32_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output | float32_t * | | |
| kernel_layout | arm_nn_dw_kernel_layout_f32 | | |
Source: `Include/Internal/arm_depthwise_conv_opt_f32.h:73`
## arm_dw_spec_k3_1d_nhwc_f32_call
`function` · `c`
```c
static arm_cmsis_nn_status arm_dw_spec_k3_1d_nhwc_f32_call(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params_f32 *params,
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *filter_dims,
const float32_t *kernel,
const cmsis_nn_dims *bias_dims,
const float32_t *bias,
const cmsis_nn_dims *output_dims,
float32_t *output,
arm_nn_dw_kernel_layout_f32 kernel_layout
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| params | const cmsis_nn_dw_conv_params_f32 * | | |
| input_dims | const cmsis_nn_dims * | | |
| input | const float32_t * | | |
| filter_dims | const cmsis_nn_dims * | | |
| kernel | const float32_t * | | |
| bias_dims | const cmsis_nn_dims * | | |
| bias | const float32_t * | | |
| output_dims | const cmsis_nn_dims * | | |
| output | float32_t * | | |
| kernel_layout | arm_nn_dw_kernel_layout_f32 | | |
Source: `Include/Internal/arm_depthwise_conv_opt_f32.h:105`
## arm_dw_spec_nhwc_f32
`attribute` · `c`
```c
const arm_dw_spec_f32 arm_dw_spec_nhwc_f32[]
```
Source: `Include/Internal/arm_depthwise_conv_opt_f32.h:135`
---
# heliaCORE.Internal.arm_minmax_f16_common
## arm_minmax_f16_impl
`function` · `c`
```c
arm_cmsis_nn_status arm_minmax_f16_impl(
const cmsis_nn_context *ctx,
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims,
int32_t select_max
)
```
Shared implementation for arm_minimum_f16 / arm_maximum_f16.
Defined out-of-line in arm_minmax_f16_common.c (rather than inline in this header) so that source-level test coverage tools (e.g. gcov) attribute hits/lines/branches to the real implementation instead of collapsing them into the call site of the thin wrapper that invokes it.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float16_t * | in | Pointer to the first input tensor data. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of the first input tensor. |
| input_2_data | const float16_t * | in | Pointer to the second input tensor data. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of the second input tensor. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| output_dims | const cmsis_nn_dims * | in | Dimensions of the output tensor. |
| select_max | int32_t | in | Non-zero to compute elementwise maximum, zero for minimum. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/Internal/arm_minmax_f16_common.h:59`
---
# heliaCORE.Internal.arm_minmax_f32_common
## arm_minmax_f32_impl
`function` · `c`
```c
arm_cmsis_nn_status arm_minmax_f32_impl(
const cmsis_nn_context *ctx,
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims,
int32_t select_max
)
```
Shared implementation for arm_minimum_f32 / arm_maximum_f32.
Defined out-of-line in arm_minmax_f32_common.c (rather than inline in this header) so that source-level test coverage tools (e.g. gcov) attribute hits/lines/branches to the real implementation instead of collapsing them into the call site of the thin wrapper that invokes it.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float32_t * | in | Pointer to the first input tensor data. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of the first input tensor. |
| input_2_data | const float32_t * | in | Pointer to the second input tensor data. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of the second input tensor. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| output_dims | const cmsis_nn_dims * | in | Dimensions of the output tensor. |
| select_max | int32_t | in | Non-zero to compute elementwise maximum, zero for minimum. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/Internal/arm_minmax_f32_common.h:59`
---
# heliaCORE.Internal.arm_nn_activation_flt
## ARM_NN_TANH_F32_XMAX
`macro` · `c`
```c
#define ARM_NN_TANH_F32_XMAX (6.0f)
```
Source: `Include/Internal/arm_nn_activation_flt.h:48`
## ARM_NN_TANH_F32_LUT_SEGMENTS
`macro` · `c`
```c
#define ARM_NN_TANH_F32_LUT_SEGMENTS (384)
```
Source: `Include/Internal/arm_nn_activation_flt.h:49`
## ARM_NN_TANH_F32_LUT_MAX_IDX
`macro` · `c`
```c
#define ARM_NN_TANH_F32_LUT_MAX_IDX (ARM_NN_TANH_F32_LUT_SEGMENTS - 1)
```
Source: `Include/Internal/arm_nn_activation_flt.h:50`
## arm_nn_hardswish_scalar_f32
`function` · `c`
```c
static float32_t arm_nn_hardswish_scalar_f32(float32_t x)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:52`
## arm_nn_tanh_scalar_ref_f32
`function` · `c`
```c
static float32_t arm_nn_tanh_scalar_ref_f32(float32_t x)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:92`
## arm_nn_sigmoid_scalar_f32
`function` · `c`
```c
static float32_t arm_nn_sigmoid_scalar_f32(float32_t x)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:177`
## arm_nn_propagate_nan_f32
`function` · `c`
```c
static float32_t arm_nn_propagate_nan_f32(float32_t x, float32_t y)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | | |
| y | float32_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:202`
## arm_nn_apply_activation_type_f32
`function` · `c`
```c
static float32_t arm_nn_apply_activation_type_f32(float32_t x, arm_nn_activation_type_flt type, float32_t act_param)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | | |
| type | arm_nn_activation_type_flt | | |
| act_param | float32_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:216`
## arm_nn_clamp_scalar_f32
`function` · `c`
```c
static float32_t arm_nn_clamp_scalar_f32(float32_t x, float32_t min_v, float32_t max_v)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | | |
| min_v | float32_t | | |
| max_v | float32_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:248`
## arm_nn_vector_clamp_f32
`function` · `c`
```c
static void arm_nn_vector_clamp_f32(
float32_t *data,
int32_t block_size,
float32_t activation_min,
float32_t activation_max
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data | float32_t * | | |
| block_size | int32_t | | |
| activation_min | float32_t | | |
| activation_max | float32_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:365`
## arm_nn_clamp_scalar_f16
`function` · `c`
```c
static float16_t arm_nn_clamp_scalar_f16(float16_t x, float16_t min_v, float16_t max_v)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | | |
| min_v | float16_t | | |
| max_v | float16_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:389`
## arm_nn_hardswish_scalar_f16
`function` · `c`
```c
static float16_t arm_nn_hardswish_scalar_f16(float16_t x)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:400`
## arm_nn_tanh_scalar_ref_f16
`function` · `c`
```c
static float16_t arm_nn_tanh_scalar_ref_f16(float16_t x)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:418`
## arm_nn_sigmoid_scalar_f16
`function` · `c`
```c
static float16_t arm_nn_sigmoid_scalar_f16(float16_t x)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:460`
## arm_nn_apply_activation_type_f16
`function` · `c`
```c
static float16_t arm_nn_apply_activation_type_f16(float16_t x, arm_nn_activation_type_flt type, float16_t act_param)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | | |
| type | arm_nn_activation_type_flt | | |
| act_param | float16_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:472`
## arm_nn_vector_clamp_f16
`function` · `c`
```c
static void arm_nn_vector_clamp_f16(
float16_t *data,
int32_t block_size,
float16_t activation_min,
float16_t activation_max
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data | float16_t * | | |
| block_size | int32_t | | |
| activation_min | float16_t | | |
| activation_max | float16_t | | |
Source: `Include/Internal/arm_nn_activation_flt.h:577`
---
# heliaCORE.Internal.arm_nn_arg_extrema_flt
## arm_nn_arg_count
`function` · `c`
```c
static bool arm_nn_arg_count(const int32_t dims, size_t width, size_t *count)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dims | const int32_t | | |
| width | size_t | | |
| count | size_t * | | |
Source: `Include/Internal/arm_nn_arg_extrema_flt.h:24`
## arm_nn_arg_load
`function` · `c`
```c
static uint32_t arm_nn_arg_load(const uint8_t *input, size_t width)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const uint8_t * | | |
| width | size_t | | |
Source: `Include/Internal/arm_nn_arg_extrema_flt.h:46`
## arm_nn_arg_key
`function` · `c`
```c
static uint32_t arm_nn_arg_key(uint32_t bits, uint32_t sign)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bits | uint32_t | | |
| sign | uint32_t | | |
Source: `Include/Internal/arm_nn_arg_extrema_flt.h:61`
## arm_nn_arg_extrema
`function` · `c`
```c
static arm_cmsis_nn_status arm_nn_arg_extrema(
const void *input_data,
const cmsis_nn_dims *input_dims,
int32_t axis,
int32_t *output_data,
size_t width,
bool maximum
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const void * | | |
| input_dims | const cmsis_nn_dims * | | |
| axis | int32_t | | |
| output_data | int32_t * | | |
| width | size_t | | |
| maximum | bool | | |
Source: `Include/Internal/arm_nn_arg_extrema_flt.h:72`
---
# heliaCORE.Internal.arm_nn_axis_copy_common
## arm_nn_axis_copy_plan
`function` · `c`
```c
static arm_cmsis_nn_status arm_nn_axis_copy_plan(
const int32_t *shape,
const int32_t dims,
const int32_t axis,
const int32_t inner_begin,
const int32_t axis_len,
const int32_t num,
const int32_t *sizes,
int32_t *outer,
int32_t *inner
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| shape | const int32_t * | | |
| dims | const int32_t | | |
| axis | const int32_t | | |
| inner_begin | const int32_t | | |
| axis_len | const int32_t | | |
| num | const int32_t | | |
| sizes | const int32_t * | | |
| outer | int32_t * | | |
| inner | int32_t * | | |
Source: `Include/Internal/arm_nn_axis_copy_common.h:45`
## arm_nn_axis_copy_ptrs_ok
`function` · `c`
```c
static int32_t arm_nn_axis_copy_ptrs_ok(const void *const *ptrs, const int32_t num, const int32_t total)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ptrs | const void *const * | | |
| num | const int32_t | | |
| total | const int32_t | | |
Source: `Include/Internal/arm_nn_axis_copy_common.h:123`
## ARM_NN_AXIS_COPY_DEFINE
`macro` · `c`
```c
#define ARM_NN_AXIS_COPY_DEFINE(SUFFIX, TYPE)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| SUFFIX | | | |
| TYPE | | | |
Source: `Include/Internal/arm_nn_axis_copy_common.h:145`
## arm_nn_axis_scatter_f32
`function` · `c`
```c
static void arm_nn_axis_scatter_f32(
const float32_t *packed,
const int32_t outer,
const int32_t num,
const int32_t *sizes,
const int32_t inner,
float32_t *const *slices
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| packed | const float32_t * | | |
| outer | const int32_t | | |
| num | const int32_t | | |
| sizes | const int32_t * | | |
| inner | const int32_t | | |
| slices | float32_t *const * | | |
Source: `Include/Internal/arm_nn_axis_copy_common.h:198`
## arm_nn_axis_gather_f32
`function` · `c`
```c
static void arm_nn_axis_gather_f32(
const float32_t *const *slices,
const int32_t outer,
const int32_t num,
const int32_t *sizes,
const int32_t inner,
float32_t *packed
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| slices | const float32_t *const * | | |
| outer | const int32_t | | |
| num | const int32_t | | |
| sizes | const int32_t * | | |
| inner | const int32_t | | |
| packed | float32_t * | | |
Source: `Include/Internal/arm_nn_axis_copy_common.h:198`
## arm_nn_axis_scatter_f16
`function` · `c`
```c
static void arm_nn_axis_scatter_f16(
const float16_t *packed,
const int32_t outer,
const int32_t num,
const int32_t *sizes,
const int32_t inner,
float16_t *const *slices
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| packed | const float16_t * | | |
| outer | const int32_t | | |
| num | const int32_t | | |
| sizes | const int32_t * | | |
| inner | const int32_t | | |
| slices | float16_t *const * | | |
Source: `Include/Internal/arm_nn_axis_copy_common.h:201`
## arm_nn_axis_gather_f16
`function` · `c`
```c
static void arm_nn_axis_gather_f16(
const float16_t *const *slices,
const int32_t outer,
const int32_t num,
const int32_t *sizes,
const int32_t inner,
float16_t *packed
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| slices | const float16_t *const * | | |
| outer | const int32_t | | |
| num | const int32_t | | |
| sizes | const int32_t * | | |
| inner | const int32_t | | |
| packed | float16_t * | | |
Source: `Include/Internal/arm_nn_axis_copy_common.h:201`
---
# heliaCORE.Internal.arm_nn_broadcast_walk
## arm_nn_broadcast_dim_valid
`function` · `c`
```c
static int32_t arm_nn_broadcast_dim_valid(const int32_t dim_1, const int32_t dim_2, const int32_t dim_out)
```
Check that one NHWC dimension of two operands broadcasts to the output dimension.
TensorFlow Lite broadcast rules: both inputs must be at least 1 (an empty tensor is rejected rather than treated as a no-op), they must be equal or one of them must be 1, and the output dimension must be the larger of the two.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dim_1 | const int32_t | | |
| dim_2 | const int32_t | | |
| dim_out | const int32_t | | |
Source: `Include/Internal/arm_nn_broadcast_walk.h:34`
## arm_nn_broadcast_dims_valid
`function` · `c`
```c
static int32_t arm_nn_broadcast_dims_valid(
const cmsis_nn_dims *dims_1,
const cmsis_nn_dims *dims_2,
const cmsis_nn_dims *dims_out
)
```
Check that two NHWC operands broadcast to the given output shape.
Every kernel that uses ARM_NN_BROADCAST_WALK_NHWC must reject arguments that fail this check, since the walk indexes each input by its own dims and writes the output by the output dims.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dims_1 | const cmsis_nn_dims * | | |
| dims_2 | const cmsis_nn_dims * | | |
| dims_out | const cmsis_nn_dims * | | |
Source: `Include/Internal/arm_nn_broadcast_walk.h:54`
## ARM_NN_BROADCAST_WALK_NHWC
`macro` · `c`
```c
#define ARM_NN_BROADCAST_WALK_NHWC(IN_TYPE, OUT_TYPE, in_1, dims_1, in_2, dims_2, out, dims_out, FULL, SCALAR_1, SCALAR_2)
```
Walk an NHWC broadcast of two operands, calling a contiguous kernel on each run.
Each input is indexed by its own dims: a dimension of 1 has stride 0 and is broadcast, any other dimension equals the output dimension and strides normally. The longest contiguous run whose shapes agree is handed to the caller's kernels, so the common cases (identical shapes, a single scalar, per-batch, per-row, per-channel) each cost one call per run.
Preconditions: arm_nn_broadcast_dims_valid(dims_1, dims_2, dims_out) is non-zero.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| IN_TYPE | | | element type of the inputs |
| OUT_TYPE | | | element type of the output |
| in_1 | | | const IN_TYPE * first input |
| dims_1 | | | const `cmsis_nn_dims` * dims of in_1 |
| in_2 | | | const IN_TYPE * second input |
| dims_2 | | | const `cmsis_nn_dims` * dims of in_2 |
| out | | | OUT_TYPE * output, sized by dims_out |
| dims_out | | | const `cmsis_nn_dims` * broadcast output dims (see arm_nn_broadcast_dims_valid) |
| FULL | | | FULL(const IN_TYPE *a, const IN_TYPE *b, OUT_TYPE *o, int32_t n) elementwise kernel over n elements of a and b |
| SCALAR_1 | | | SCALAR_1(const IN_TYPE *scalar, const IN_TYPE *vec, OUT_TYPE *o, int32_t n) kernel where *scalar is one element of in_1 broadcast against n elements of in_2 |
| SCALAR_2 | | | SCALAR_2(const IN_TYPE *scalar, const IN_TYPE *vec, OUT_TYPE *o, int32_t n) kernel where *scalar is one element of in_2 broadcast against n elements of in_1; note the operands arrive in reversed order, so an asymmetric kernel must swap them |
Source: `Include/Internal/arm_nn_broadcast_walk.h:89`
---
# heliaCORE.Internal.arm_nn_config
## ARM_NN_FLOAT_API_ENABLED
`macro` · `c`
```c
#define ARM_NN_FLOAT_API_ENABLED (ARM_NN_ENABLE_F32 || ARM_NN_ENABLE_F16)
```
Optional feature gates for floating-point extensions.
These are disabled by default so integer-only builds keep their current code size and API surface. The float16 feature gate still depends on toolchain and target support such as ARM_FLOAT16_SUPPORTED where applicable.
Source: `Include/Internal/arm_nn_config.h:62`
---
# heliaCORE.Internal.arm_nn_pool_window_common
## arm_nn_pool_axis_valid
`function` · `c`
```c
static bool arm_nn_pool_axis_valid(const int32_t n, const int32_t s, const int32_t p, const int32_t k, const int32_t w)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| n | const int32_t | | |
| s | const int32_t | | |
| p | const int32_t | | |
| k | const int32_t | | |
| w | const int32_t | | |
Source: `Include/Internal/arm_nn_pool_window_common.h:40`
---
# heliaCORE.Internal.arm_nn_s4_decode
## arm_nn_s4_low_nibble
`function` · `c`
```c
static int8_t arm_nn_s4_low_nibble(int8_t packed)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| packed | int8_t | | |
Source: `Include/Internal/arm_nn_s4_decode.h:17`
## arm_nn_s4_high_nibble
`function` · `c`
```c
static int8_t arm_nn_s4_high_nibble(int8_t packed)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| packed | int8_t | | |
Source: `Include/Internal/arm_nn_s4_decode.h:25`
---
# heliaCORE.Internal.arm_nn_sqrt_flt
## ARM_NN_SQRT_EXACT_FN
`macro` · `c`
```c
#define ARM_NN_SQRT_EXACT_FN
```
Source: `Include/Internal/arm_nn_sqrt_flt.h:53`
## arm_nn_f32_to_bits
`function` · `c`
```c
static uint32_t arm_nn_f32_to_bits(float32_t value)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| value | float32_t | | |
Source: `Include/Internal/arm_nn_sqrt_flt.h:58`
## arm_nn_f32_from_bits
`function` · `c`
```c
static float32_t arm_nn_f32_from_bits(uint32_t bits)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bits | uint32_t | | |
Source: `Include/Internal/arm_nn_sqrt_flt.h:65`
## arm_nn_sqrt_special_f32
`function` · `c`
```c
static bool arm_nn_sqrt_special_f32(uint32_t in_bits, bool reciprocal, uint32_t *out_bits)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_bits | uint32_t | | |
| reciprocal | bool | | |
| out_bits | uint32_t * | | |
Source: `Include/Internal/arm_nn_sqrt_flt.h:73`
## arm_nn_f16_to_bits
`function` · `c`
```c
static uint16_t arm_nn_f16_to_bits(float16_t value)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| value | float16_t | | |
Source: `Include/Internal/arm_nn_sqrt_flt.h:104`
## arm_nn_f16_from_bits
`function` · `c`
```c
static float16_t arm_nn_f16_from_bits(uint16_t bits)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bits | uint16_t | | |
Source: `Include/Internal/arm_nn_sqrt_flt.h:111`
## arm_nn_sqrt_special_f16
`function` · `c`
```c
static bool arm_nn_sqrt_special_f16(uint16_t in_bits, bool reciprocal, uint16_t *out_bits)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_bits | uint16_t | | |
| reciprocal | bool | | |
| out_bits | uint16_t * | | |
Source: `Include/Internal/arm_nn_sqrt_flt.h:119`
---
# heliaCORE.Internal.arm_resize_nearest_neighbor_common
## arm_nn_resize_nearest_neighbor_scratch_bytes
`function` · `c`
```c
static int32_t arm_nn_resize_nearest_neighbor_scratch_bytes(const cmsis_nn_dims *output_dims)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| output_dims | const cmsis_nn_dims * | | |
Source: `Include/Internal/arm_resize_nearest_neighbor_common.h:39`
## arm_nn_resize_nearest_neighbor_prepare
`function` · `c`
```c
static arm_cmsis_nn_status arm_nn_resize_nearest_neighbor_prepare(
const cmsis_nn_context *ctx,
const cmsis_nn_resize_params *resize_params,
const cmsis_nn_dims *input_shape,
const cmsis_nn_dims *output_size_shape,
const int32_t *output_size_data,
const cmsis_nn_dims *output_shape,
int32_t **x_map_out,
int32_t **y_map_out
)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | | |
| resize_params | const cmsis_nn_resize_params * | | |
| input_shape | const cmsis_nn_dims * | | |
| output_size_shape | const cmsis_nn_dims * | | |
| output_size_data | const int32_t * | | |
| output_shape | const cmsis_nn_dims * | | |
| x_map_out | int32_t ** | | |
| y_map_out | int32_t ** | | |
Source: `Include/Internal/arm_resize_nearest_neighbor_common.h:50`
## ARM_RESIZE_NEAREST_NEIGHBOR_DEFINE
`macro` · `c`
```c
#define ARM_RESIZE_NEAREST_NEIGHBOR_DEFINE(FUNC_NAME, SCALAR_T, MEMCPY_FUNC)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| FUNC_NAME | | | |
| SCALAR_T | | | |
| MEMCPY_FUNC | | | |
Source: `Include/Internal/arm_resize_nearest_neighbor_common.h:129`
---
# heliaCORE.Internal.arm_strided_slice_common
## ARM_STRIDED_SLICE_DEFINE
`macro` · `c`
```c
#define ARM_STRIDED_SLICE_DEFINE(FUNC_NAME, SCALAR_T, MEMCPY_FUNC)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| FUNC_NAME | | | |
| SCALAR_T | | | |
| MEMCPY_FUNC | | | |
Source: `Include/Internal/arm_strided_slice_common.h:32`
---
# heliaCORE.Internal.arm_transpose_common
## ARM_TRANSPOSE_DEFINE
`macro` · `c`
```c
#define ARM_TRANSPOSE_DEFINE(FUNC_NAME, SCALAR_T, PARAMS_T, LAYOUT_T, TRANSPOSE_2D_FUNC)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| FUNC_NAME | | | |
| SCALAR_T | | | |
| PARAMS_T | | | |
| LAYOUT_T | | | |
| TRANSPOSE_2D_FUNC | | | |
Source: `Include/Internal/arm_transpose_common.h:35`
---
# heliaCORE.arm_nn_math_types_flt
## float32_t
`type` · `c`
```c
typedef float float32_t
```
32-bit floating-point type definition for CMSIS-NN float extensions.
Source: `Include/arm_nn_math_types_flt.h:45`
## ARM_NN_F32_FINITE_MAX
`macro` · `c`
```c
#define ARM_NN_F32_FINITE_MAX ((float32_t)__FLT_MAX__)
```
Largest finite float32 value representable by the toolchain.
Source: `Include/arm_nn_math_types_flt.h:52`
## ARM_NN_F32_FINITE_LOWEST
`macro` · `c`
```c
#define ARM_NN_F32_FINITE_LOWEST (-ARM_NN_F32_FINITE_MAX)
```
Lowest finite float32 value representable by the toolchain.
Source: `Include/arm_nn_math_types_flt.h:57`
## ARM_NN_F16_FINITE_MAX
`macro` · `c`
```c
#define ARM_NN_F16_FINITE_MAX ((float16_t)65504.0f)
```
Largest finite float16 value representable by the toolchain.
Source: `Include/arm_nn_math_types_flt.h:116`
## ARM_NN_F16_FINITE_LOWEST
`macro` · `c`
```c
#define ARM_NN_F16_FINITE_LOWEST ((float16_t)(-65504.0f))
```
Lowest finite float16 value representable by the toolchain.
Negated before the cast: negating a float16_t that is __fp16 promotes it to float, which -Wdouble-promotion reports at every use.
Source: `Include/arm_nn_math_types_flt.h:128`
---
# heliaCORE.arm_nn_math_types
## NN_Q31_MAX
`macro` · `c`
```c
#define NN_Q31_MAX ((int32_t)(0x7FFFFFFFL))
```
Translate architecture feature flags to CMSIS-NN defines.
Limits macros
Source: `Include/arm_nn_math_types.h:117`
## NN_Q15_MAX
`macro` · `c`
```c
#define NN_Q15_MAX ((int16_t)(0x7FFF))
```
Source: `Include/arm_nn_math_types.h:118`
## NN_Q7_MAX
`macro` · `c`
```c
#define NN_Q7_MAX ((int8_t)(0x7F))
```
Source: `Include/arm_nn_math_types.h:119`
## NN_Q31_MIN
`macro` · `c`
```c
#define NN_Q31_MIN ((int32_t)(0x80000000L))
```
Source: `Include/arm_nn_math_types.h:120`
## NN_Q15_MIN
`macro` · `c`
```c
#define NN_Q15_MIN ((int16_t)(0x8000))
```
Source: `Include/arm_nn_math_types.h:121`
## NN_Q7_MIN
`macro` · `c`
```c
#define NN_Q7_MIN ((int8_t)(0x80))
```
Source: `Include/arm_nn_math_types.h:122`
---
# heliaCORE.arm_nn_tables
## sigmoid_table_uint16
`attribute` · `c`
```c
const uint16_t sigmoid_table_uint16[256]
```
tables for various activation functions
Source: `Include/arm_nn_tables.h:40`
---
# heliaCORE.arm_nn_types
## NS_CMSIS_NN_VERSION_MAJOR
`macro` · `c`
```c
#define NS_CMSIS_NN_VERSION_MAJOR (7) /* x-release-please-major */
```
Source: `Include/arm_nn_types.h:41`
## NS_CMSIS_NN_VERSION_MINOR
`macro` · `c`
```c
#define NS_CMSIS_NN_VERSION_MINOR (39) /* x-release-please-minor */
```
Source: `Include/arm_nn_types.h:42`
## NS_CMSIS_NN_VERSION_PATCH
`macro` · `c`
```c
#define NS_CMSIS_NN_VERSION_PATCH (2) /* x-release-please-patch */
```
Source: `Include/arm_nn_types.h:43`
## NS_CMSIS_NN
`macro` · `c`
```c
#define NS_CMSIS_NN (1)
```
Identity macros for the ns-cmsis-nn (Ambiq) superset of CMSIS-NN.
This library is wire-compatible with upstream ARM-software/CMSIS-NN: every upstream `arm_*` symbol resolves here. We additionally ship Ambiq-specific kernels (e.g. arm_gather_s8, the elementwise prelu/clamp variants, ...).
Downstream code that depends on Ambiq-only kernels should guard with:
```
#if !defined(NS_CMSIS_NN)
# error "this code requires ns-cmsis-nn (Ambiq superset)"
#endif
#if NS_CMSIS_NN_VERSION < 7024000
# error "needs ns-cmsis-nn >= 7.24.0"
#endif
```
NS_CMSIS_NN_VERSION is packed as MAJOR * 1000000 + MINOR * 1000 + PATCH (each of MINOR and PATCH gets a full 3-digit field, so semantic ordering is preserved for any reasonable version). It tracks release-please bumps automatically through the per-component macros above; no separate marker is required.
Source: `Include/arm_nn_types.h:67`
## NS_CMSIS_NN_VERSION
`macro` · `c`
```c
#define NS_CMSIS_NN_VERSION ((NS_CMSIS_NN_VERSION_MAJOR * 1000000) + (NS_CMSIS_NN_VERSION_MINOR * 1000) + NS_CMSIS_NN_VERSION_PATCH)
```
Source: `Include/arm_nn_types.h:68`
---
# heliaCORE.arm_nnfunctions
## USE_INTRINSIC
`macro` · `c`
```c
#define USE_INTRINSIC
```
Source: `Include/arm_nnfunctions.h:47`
## arm_convolve_wrapper_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_wrapper_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_wrapper_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 packed with 2x int4 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:97`
## arm_convolve_wrapper_s4_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_wrapper_s4_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required buffer size for arm_convolve_wrapper_s4.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:133`
## arm_convolve_wrapper_s4_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_convolve_wrapper_s4_get_buffer_size_mve(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required buffer size for arm_convolve_wrapper_s4 for Arm(R) Helium Architecture case.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s4_get_buffer_size()`. Currently this operator does not have an mve implementation, so dsp will be used.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a 0 from a route needing no buffer is not a statement that the shape is valid.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:149`
## arm_convolve_wrapper_s4_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_convolve_wrapper_s4_get_buffer_size_dsp(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required buffer size for arm_convolve_wrapper_s4 for processors with DSP extension.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s4_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a 0 from a route needing no buffer is not a statement that the shape is valid.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:164`
## arm_convolve_wrapper_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_wrapper_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
- On builds with ARM_MATH_MVEI (without ARM_MATH_AUTOVECTORIZE), a layer that would run `arm_convolve_s8()` and is in the gate of `arm_convolve_s8_small_cin()` or `arm_convolve_s8_3x3_c16_s1()` runs that entry instead, with the same result, scratch and weight sums. The input depth is checked first, so other layers skip both gates.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_wrapper_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. The selected kernel only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build: this wrapper dispatches to `arm_convolve_s8()`, `arm_convolve_1x1_s8()`, `arm_convolve_1x1_s8_fast()`, `arm_convolve_1_x_n_s8()` and `arm_convolve_1x1_out_s8()`. The buffer contents are consumed only on builds with the MVE extension (ARM_MATH_MVEI), and on those builds every one of those kernels diagnoses a NULL buf with ARM_CMSIS_NN_ARG_ERROR; on other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed that way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:225`
## arm_convolve_wrapper_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_wrapper_s8_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required buffer size for arm_convolve_wrapper_s8.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:264`
## arm_convolve_s8_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_convolve_s8_get_buffer_size_mve(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for arm_convolve_s8 for Arm(R) Helium Architecture case.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas, so a shape whose dimension product overflows is only reported by the leg that actually computes a buffer for it.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_convolve_s8_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:278`
## arm_convolve_wrapper_s8_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_convolve_wrapper_s8_get_buffer_size_mve(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required buffer size for arm_convolve_wrapper_s8 for Arm(R) Helium Architecture case.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s8_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a 0 from a route needing no buffer is not a statement that the shape is valid.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:290`
## arm_convolve_wrapper_s8_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_convolve_wrapper_s8_get_buffer_size_dsp(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required buffer size for arm_convolve_wrapper_s8 for processors with DSP extension.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s8_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a 0 from a route needing no buffer is not a statement that the shape is valid.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:305`
## arm_convolve_wrapper_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_wrapper_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_wrapper_s8_get_buffer_size will return the buffer_size if required The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const cmsis_nn_bias_data * | in | Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:338`
## arm_convolve_s16_group_ch_mult_1
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_s16_group_ch_mult_1(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
s16 grouped convolution optimized for the case where filter_dims->c == 1 and input_ch == output_ch (channel multiplier = 1).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context (unused, pass NULL-initialised). |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (strides, dilations, pads, activation). |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info (multiplier and shift). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int16_t * | in | Input data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, 1] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions (unused, may be zero-initialised). |
| bias_data | const cmsis_nn_bias_data * | in | Optional bias struct (int32 or int64). May be NULL. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success. |
Source: `Include/arm_nnfunctions.h:368`
## arm_convolve_wrapper_s16_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_wrapper_s16_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required buffer size for arm_convolve_wrapper_s16.
An out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:398`
## arm_convolve_wrapper_s16_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_convolve_wrapper_s16_get_buffer_size_dsp(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required buffer size for arm_convolve_wrapper_s16 for for processors with DSP extension.
An out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s16_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:412`
## arm_convolve_wrapper_s16_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_convolve_wrapper_s16_get_buffer_size_mve(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required buffer size for arm_convolve_wrapper_s16 for Arm(R) Helium Architecture case.
An out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s16_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:426`
## arm_convolve_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Basic s4 convolution function.
1. Supported framework: TensorFlow Lite micro
2. Additional memory is required for optimization. Refer to argument 'ctx' for details.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| filter_data | const int8_t * | in | Packed Filter data pointer. Data type: int8 packed with 2x int4 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:458`
## arm_convolve_even_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_even_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Basic s4 convolution function with a requirement of even number of kernels.
1. Supported framework: TensorFlow Lite micro
2. Additional memory is required for optimization. Refer to argument 'ctx' for details.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_even_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions. Note the product must be even. |
| filter_data | const int8_t * | in | Packed Filter data pointer. Data type: int8 packed with 2x int4 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` if not for MVE |
Source: `Include/arm_nnfunctions.h:499`
## arm_convolve_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *upscale_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Basic s8 convolution function.
1. Supported framework: TensorFlow Lite micro
2. Additional memory is required for optimization. Refer to argument 'ctx' for details.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. For grouped convolution the entries run over all output_dims->c channels, groups laid out consecutively. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI); an unfilled buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, and a NULL buf is currently reported as ARM_CMSIS_NN_ARG_ERROR. On other builds this function currently derives the same quantity itself and does not read the context. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] where HK, WK and CK are the spatial filter dimensions. CK != C_IN is used for grouped convolution, in which case the required conditions are C_IN = N * CK and C_OUT = N * M for N groups of size M. |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions for transpose. Format: [H_UP, W_UP] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` |
Source: `Include/arm_nnfunctions.h:563`
## arm_convolve_s8_small_cin
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_s8_small_cin(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *upscale_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model. It copies each kernel row with one predicated vector load and multiplies four output channels per step.
- The output is identical to `arm_convolve_s8()`. The bias is read through the weight sums, which `arm_convolve_weight_sum()` fills as for `arm_convolve_s8()`; bias_dims and bias_data are unused.
- Gate: upscale_dims NULL, C_IN from 1 to 3 with CK equal to C_IN (one group), dilation 1 in both dimensions, WK and HK at least 1 with WK x C_IN at most 16 and HK x WK x C_IN at most 48, and C_OUT a positive multiple of 4. Stride, padding and batch count are as for `arm_convolve_s8()`.
- Scratch: ctx->buf holds `arm_convolve_s8_get_buffer_size()` bytes (4 x 16 x ceil(HK x WK x C_IN / 16) on ARM_MATH_MVEI builds), the same as `arm_convolve_s8()`, and needs no alignment.
- It is a direct entry: `arm_convolve_s8()` does not call it, and `arm_convolve_wrapper_s8()` calls it for layers in the gate that it would otherwise pass to `arm_convolve_s8()`. A caller that selects the kernel per layer ahead of time calls it for layers in the gate and `arm_convolve_s8()` for every other layer, or on `ARM_CMSIS_NN_NO_IMPL_ERROR`. Both take the same arguments, scratch and weight sums.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context with `arm_convolve_s8_get_buffer_size()` bytes of scratch, all of which may be written |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, as for `arm_convolve_s8()` |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters, as for `arm_convolve_s8()` |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions for transpose. Format: [H_UP, W_UP] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - an argument error that `arm_convolve_s8()` reports: ctx->buf is NULL, C_IN or C_OUT is not a multiple of the group count C_IN / CK, or weight_sum_ctx->buf is NULL on builds with ARM_MATH_MVEI. These are checked before the gate. `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer is outside the gate below, or the build lacks ARM_MATH_MVEI or defines ARM_MATH_AUTOVECTORIZE; nothing is written, to the output or to the scratch `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:620`
## arm_convolve_s8_3x3_c16_s1
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_s8_3x3_c16_s1(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *upscale_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 3x3 convolution over 16 input channels with unit stride. It reads the kernel rows of a patch inside the input in place, copying only patches that cross the border, and multiplies four output pixels per filter load.
- The output is identical to `arm_convolve_s8()`. The bias is read through the weight sums; bias_dims and bias_data are unused.
- Gate: upscale_dims NULL, C_IN and CK both 16 (one group), HK and WK both 3, and stride and dilation 1 in both dimensions. Padding, batch count and C_OUT are as for `arm_convolve_s8()`.
- It is a direct entry: `arm_convolve_s8()` does not call it, and `arm_convolve_wrapper_s8()` calls it for layers in the gate that it would otherwise pass to `arm_convolve_s8()`. A caller that selects the kernel per layer ahead of time calls it for layers in the gate and `arm_convolve_s8()` for every other layer, or on `ARM_CMSIS_NN_NO_IMPL_ERROR`. The gate does not overlap that of `arm_convolve_s8_small_cin()`.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context with `arm_convolve_s8_get_buffer_size()` bytes of scratch (576 bytes on ARM_MATH_MVEI builds), all of which may be written |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, as for `arm_convolve_s8()` |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters, as for `arm_convolve_s8()` |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions for transpose. Format: [H_UP, W_UP] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_convolve_s8_small_cin()` `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer is outside the gate below, or the build lacks ARM_MATH_MVEI or defines ARM_MATH_AUTOVECTORIZE; nothing is written, to the output or to the scratch `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:672`
## arm_convolve_s4_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_s4_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for s4 convolution function.
The dimensions and the byte count are both checked here, so an out-of-range shape returns -1 on every build target rather than a wrapped size.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:698`
## arm_convolve_even_s4_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_even_s4_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for arm_convolve_even_s4.
Forwards to `arm_convolve_s4_get_buffer_size()`: the even_s4 kernel stages up to four im2col rows of filter_dims->w * filter_dims->h * input_dims->c int8 elements, byte-for-byte the size that sizer returns. The equality, including the -1 answers for out-of-range shapes, is pinned by a Unity test.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:713`
## arm_convolve_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_s8_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for s8 convolution function.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas, so a shape whose dimension product overflows is only reported by the leg that actually computes a buffer for it.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:729`
## arm_convolve_s8_get_weights_sum_size
`function` · `c`
```c
int32_t arm_convolve_s8_get_weights_sum_size(const cmsis_nn_dims *output_dims)
```
Get the required buffer size for s8 convolution and depthwise convolution weight sum.
For a valid (non-negative, in-range) output_dims->c, returns output_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. A negative or out-of-range output_dims->c returns -1 on builds with the MVE extension; elsewhere no weight sum buffer is used and the answer stays 0.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| output_dims | const cmsis_nn_dims * | in | Output (activation) tensor dimensions. Format: [N, H, W, C_COUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required weight sum buffer size in bytes, or -1 if output_dims->c is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:742`
## arm_transpose_conv_wrapper_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_conv_wrapper_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_context *reverse_conv_ctx,
const cmsis_nn_transpose_conv_params *transpose_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Wrapper to select optimal transposed convolution algorithm depending on parameters.
1. Supported framework: TensorFlow Lite micro
2. Additional memory is required for optimization. Refer to arguments 'ctx' and 'reverse_conv_ctx' for details.
3. Dilation is not supported: transpose_conv_params->dilation must be 1 in both dimensions, otherwise `ARM_CMSIS_NN_ARG_ERROR` is returned.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. `arm_transpose_conv_s8_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. The function only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and transpose_conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing transpose_conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Compute the sums over filter_data exactly as passed to this function: this wrapper guarantees that whatever filter preparation it performs internally preserves the per-output-channel sums, so no reversed or otherwise rearranged copy of the weights is needed for this step. Pass a valid context on every build. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the reverse-convolution route forwards this context to `arm_convolve_s8()`; an unfilled buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, and a NULL buf is currently reported as ARM_CMSIS_NN_ARG_ERROR. On other builds the contents are currently not read. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
| reverse_conv_ctx | const cmsis_nn_context * | in, out | Function context for the reversed filter used when this wrapper routes to the reverse convolution. Holds filter height * filter width * input channels * output channels int8 values; `arm_transpose_conv_s8_get_reverse_conv_buffer_size()` returns the required size (0 when the reverse-convolution route is not taken). The caller is expected to clear the buffer, if applicable, for security reasons. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of transpose_conv_params->input_offset : [-127, 128] Range of transpose_conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each out channel. |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:812`
## arm_transpose_conv_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_conv_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_transpose_conv_params *transpose_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Basic s8 transpose convolution function.
1. Supported framework: TensorFlow Lite micro
2. Additional memory is required for optimization. Refer to argument 'ctx' for details; 'output_ctx' is unused.
3. Dilation is not supported: transpose_conv_params->dilation must be 1 in both dimensions, otherwise `ARM_CMSIS_NN_ARG_ERROR` is returned.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_transpose_conv_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| output_ctx | const cmsis_nn_context * | in, out | Not accessed by this function: its buffer is neither read nor written, and it therefore has no size requirement. The parameter exists only to keep one signature across the transpose-conv family, whose float twins ignore it the same way; `arm_transpose_conv_wrapper_s8()` forwards its reverse_conv_ctx into this slot. In-tree callers pass a valid context, whose buf may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of transpose_conv_params->input_offset : [-127, 128] Range of transpose_conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each out channel. |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:864`
## arm_transpose_conv_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_transpose_conv_s8_get_buffer_size(
const cmsis_nn_transpose_conv_params *transposed_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *out_dims
)
```
Get the required buffer size for ctx in s8 transpose conv function.
The returned size is safe for both `arm_transpose_conv_s8()` and `arm_transpose_conv_wrapper_s8()`: it is the larger of the two routes' requirements, so it may exceed what the wrapper's reverse-convolution route alone would need. When either route is out of range the sentinel is propagated ahead of that comparison, so -1 is never collapsed into a plausible positive size by the other route.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transposed_conv_params | const cmsis_nn_transpose_conv_params * | in | Transposed convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| out_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative, either stride is not positive, or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:894`
## arm_transpose_conv_s8_get_reverse_conv_buffer_size
`function` · `c`
```c
int32_t arm_transpose_conv_s8_get_reverse_conv_buffer_size(
const cmsis_nn_transpose_conv_params *transposed_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims
)
```
Get the required buffer size for output_ctx in s8 transpose conv function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transposed_conv_params | const cmsis_nn_transpose_conv_params * | in | Transposed convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:910`
## arm_transpose_conv_s8_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_transpose_conv_s8_get_buffer_size_mve(
const cmsis_nn_transpose_conv_params *transposed_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *out_dims
)
```
Get size of additional buffer required by `arm_transpose_conv_s8()` for Arm(R) Helium Architecture case.
The returned size is safe for both `arm_transpose_conv_s8()` and `arm_transpose_conv_wrapper_s8()`: it is the larger of the two routes' requirements, so it may exceed what the wrapper's reverse-convolution route alone would need. When either route is out of range the sentinel is propagated ahead of that comparison, so -1 is never collapsed into a plausible positive size by the other route.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_transpose_conv_s8_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transposed_conv_params | const cmsis_nn_transpose_conv_params * | in | Transposed convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| out_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative, either stride is not positive, or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:923`
## arm_convolve_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
Basic s16 convolution function.
1. Supported framework: TensorFlow Lite micro
2. Additional memory is required for optimization. Refer to argument 'ctx' for details.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const cmsis_nn_bias_data * | in | Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` |
Source: `Include/arm_nnfunctions.h:958`
## arm_convolve_1x1_s16_ns_np_nd
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_s16_ns_np_nd(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
Pointwise s16 convolution function: no stride, no padding, no dilation.
1. Supported framework: TensorFlow Lite micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const cmsis_nn_bias_data * | in | Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` |
Source: `Include/arm_nnfunctions.h:999`
## arm_convolve_s16_fast_small_kernel
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_s16_fast_small_kernel(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
arm_convolve_s16_fast_small_kernel function. The kernel size is <=8
1. Supported framework: TensorFlow Lite micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const cmsis_nn_bias_data * | in | Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` |
Source: `Include/arm_nnfunctions.h:1040`
## arm_convolve_s16_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_s16_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for s16 convolution function.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:1065`
## arm_convolve_1x1_s4_fast
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_s4_fast(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Fast s4 version for 1x1 convolution (non-square shape).
- Supported framework : TensorFlow Lite Micro
- The following constrains on the arguments apply
1. conv_params->padding.w = conv_params->padding.h = 0
2. conv_params->stride.w = conv_params->stride.h = 1
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_1x1_s4_fast_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 packed with 2x int4 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:1098`
## arm_convolve_1x1_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s4 version for 1x1 convolution with support for non-unity stride values
- Supported framework : TensorFlow Lite Micro
- The following constrains on the arguments apply
1. conv_params->padding.w = conv_params->padding.h = 0
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. None is required by this function. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 packed with 2x int4 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:1138`
## arm_convolve_1x1_s8_fast
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_s8_fast(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Fast s8 version for 1x1 convolution (non-square shape).
- Supported framework : TensorFlow Lite Micro
- The following constrains on the arguments apply
1. conv_params->padding.w = conv_params->padding.h = 0
2. conv_params->stride.w = conv_params->stride.h = 1
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_1x1_s8_fast_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. This function reads the buffer contents only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. On other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. Note also that on an Arm Compiler build (__ARMCC_VERSION >= 6010050) with ARM_MATH_DSP and without ARM_MATH_MVEI, supplying ctx->buf selects a buffered path that never reads weight_sum_ctx. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:1204`
## arm_convolve_1x1_s4_fast_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_1x1_s4_fast_get_buffer_size(const cmsis_nn_dims *input_dims)
```
Get the required buffer size for arm_convolve_1x1_s4_fast.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns the required buffer size in bytes, or -1 if input_dims->c is negative. No build needs this scratch buffer, so every valid shape returns 0. |
Source: `Include/arm_nnfunctions.h:1225`
## arm_convolve_1x1_s8_fast_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_1x1_s8_fast_get_buffer_size(const cmsis_nn_dims *input_dims)
```
Get the required buffer size for arm_convolve_1x1_s8_fast.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns the required buffer size in bytes, or -1 if input_dims->c is negative. On builds that need this scratch buffer it also returns -1 if the required size would not fit in an int32_t; other builds need no buffer and return 0. |
Source: `Include/arm_nnfunctions.h:1236`
## arm_convolve_1x1_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 version for 1x1 convolution with support for non-unity stride values
- Supported framework : TensorFlow Lite Micro
- The following constrains on the arguments apply
1. conv_params->padding.w = conv_params->padding.h = 0
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. None is required by this function. |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. This function reads the buffer contents only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. On other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:1285`
## arm_convolve_1_x_n_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1_x_n_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
1xn convolution
- Supported framework : TensorFlow Lite Micro
- The following constraints on the arguments apply
1. input_dims->h, filter_dims->h and output_dims->h equal 1, and conv_params->padding.h is 0
2. conv_params->dilation.w is 1 and conv_params->stride.w is positive
3. conv_params->stride.w * input_dims->c is a multiple of 4
4. conv_params->padding.w, input_dims->w and output_dims->w are not negative, and filter_dims->w is at least 1
- Any horizontal padding and output width are handled, including an odd total padding, a filter wider than the input and a VALID layer whose stride leaves trailing input unused. On MVE builds the output columns whose window starts before or ends past the input read a padded copy of the input columns they span, staged in ctx; the other columns read the input in place.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_1_x_n_s8_get_buffer_size will return the buffer_size if required. buf must not be NULL. On builds with the MVE extension (ARM_MATH_MVEI) a non-zero ctx->size smaller than the staging the layer needs is rejected with ARM_CMSIS_NN_ARG_ERROR. The caller is expected to clear the buffer, if applicable, for security reasons. |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. The contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR; on other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:1357`
## arm_convolve_weight_sum
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_weight_sum(
int32_t *vector_sum_buf,
const int8_t *rhs,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
const int32_t lhs_offset,
const int32_t *bias_data
)
```
Pre-computes per-output-channel weight sums for a standard convolution.
- Supported framework : TensorFlow Lite Micro
- The buffer pointed to by `vector_sum_buf` must be at least `output_dims->c × sizeof(int32_t)` bytes. `arm_convolve_s8_get_weights_sum_size()` returns that size on builds that use the sums, 0 elsewhere, and -1 for an output_dims->c that is negative or too large to size.
- Layout: one int32 per output channel, indexed 0..`output_dims->c - 1`. Entry j holds `lhs_offset * sum(weights of output channel j) + bias_data[j]`, i.e. the bias and the input-offset contribution folded together. For grouped convolution the entries run over all output channels, with the groups laid out consecutively.
- This is the buffer the `weight_sum_ctx` parameter of the s8 convolution kernels carries. Those kernels currently treat it as an input they only read, so it has to be filled before the call - see the individual functions for what each one currently does on MVE and non-MVE builds.
- Reuse and invalidation: the contents depend only on `rhs`, `bias_data` and `lhs_offset`. They do not depend on the activations, so a buffer stays valid across calls and across batches for as long as those three are unchanged - for a static model the sums can be computed once at load time rather than per inference. Recompute whenever the weights, the bias or the input offset change (for example on requantization or a weight reload). The buffer is sized by one layer's `output_dims->c` and is specific to that layer's weights, so it cannot be shared between layers; give each layer its own.
- Returns `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| vector_sum_buf | int32_t * | out | Pointer to the buffer that will hold the weight sums. |
| rhs | const int8_t * | in | Pointer to the filter weights. Data type: int8 |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, KH, KW, C_IN] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| lhs_offset | const int32_t | in | Input-offset added to every input element before MAC. Range: [-127, 128] |
| bias_data | const int32_t * | in | Optional bias pointer. Data type: int32 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments, `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed and the buffer is left untouched, or `ARM_CMSIS_NN_SUCCESS` on success. Portable code should not treat the `NO_IMPL_ERROR` case as a failure. |
Source: `Include/arm_nnfunctions.h:1409`
## arm_depthwise_convolve_weight_sum
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_convolve_weight_sum(
int32_t *vector_sum_buf,
int8_t *scratch_buf,
const int8_t *rhs,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
const int32_t lhs_offset,
const int32_t *bias_data
)
```
Pre-computes per-channel weight sums for a depthwise convolution.
- Supported framework : TensorFlow Lite Micro
- Layout: one int32 per channel, sized by `arm_convolve_s8_get_weights_sum_size()`. Entry j holds `bias_data[j] + lhs_offset * sum(kernel values of channel j)`.
- Reuse and invalidation follow the same rules as `arm_convolve_weight_sum()`: the contents depend only on `rhs`, `bias_data` and `lhs_offset`, so they may be computed once and reused until one of those changes, and they are specific to a single layer.
- Returns `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed.
- Not interchangeable with `arm_convolve_weight_sum()`: this function walks the channel-interleaved depthwise layout `[1, KH, KW, C_OUT]` with a stride of C_OUT, whereas `arm_convolve_weight_sum()` sums contiguous runs of `KH * KW * C_IN` weights. The two agree only by coincidence. Several in-tree tests do fill a depthwise weight_sum_ctx with `arm_convolve_weight_sum()` and are still correct, for one of three unrelated reasons: `arm_depthwise_conv_wrapper_s8()` does not consume the buffer on that route at all (ch_mult != 1, batches != 1, or a dilation the optimized route does not take - see that function); the wrapper converts the layer to a regular convolution, so conv-style sums are what is wanted; or C_OUT is 1, which collapses the stride-C_OUT walk to a contiguous one and makes the two helpers compute identical values. None of those generalise, so do not read them as licence to substitute one helper for the other. Use this function wherever the sums are actually read.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| vector_sum_buf | int32_t * | out | Buffer to hold the computed weight sums. |
| scratch_buf | int8_t * | in, out | Currently unused: the implementation does not read or write it on any build, so NULL is accepted. Retained for signature compatibility; if a real buffer is passed, the caller is expected to clear it for security reasons. |
| rhs | const int8_t * | in | Depthwise convolution weights. Data type: int8 |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise-convolution parameters (stride, dilation, pad, etc.) |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, KH, KW, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| lhs_offset | const int32_t | in | Input-offset applied before MAC. Range: [-127, 128] |
| bias_data | const int32_t * | in | Optional bias pointer. Data type: int32 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments, `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed and the buffer is left untouched, or `ARM_CMSIS_NN_SUCCESS` on success. Portable code should not treat the `NO_IMPL_ERROR` case as a failure. |
Source: `Include/arm_nnfunctions.h:1458`
## arm_convolve_1x1_out_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1x1_out_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Optimised convolution for 1x1 output images (shape of BX1x1xC_OUT) for 8x8 computations.
- Supported framework : TensorFlow Lite Micro
- Optimised for Bx1×1xC output CNN layers.
- Constraints:
1. `output_dims->h` and `output_dims->w` must equal 1
2. `output_dims->c` is expected to be a multiple of 4 for best performance
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that supplies a scratch buffer for activation rearrangement. A NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. The buffer must hold one 4-byte-aligned GEMM row, that is round_up_4(filter_dims->h * filter_dims->w * filter_dims->c) bytes, as returned by `arm_convolve_1x1_out_s8_get_buffer_size()`. The requirement does not scale with the group count: the kernel rewinds its im2col cursor to the start of the buffer after each group. Setting ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check, which is what TFLite Micro and derivatives do today. The caller is expected to clear the buffer, if applicable, for security reasons. |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. The contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR; on other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (stride, dilation, pad, offsets). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantisation multipliers and shifts. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Pointer to input data. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, KH, KW, C_IN] |
| filter_data | const int8_t * | in | Pointer to filter data. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, 1, 1, C_OUT] |
| output_data | int8_t * | out | Pointer to output data. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_ARG_ERROR` on bad args, or `ARM_CMSIS_NN_SUCCESS` on success. |
Source: `Include/arm_nnfunctions.h:1523`
## arm_convolve_1x1_out_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_1x1_out_s8_get_buffer_size(const cmsis_nn_dims *filter_dims)
```
Get the required scratch buffer size for `arm_convolve_1x1_out_s8()`.
:::note
The figure is independent of the group count. `arm_convolve_1x1_out_s8()` rewinds its im2col cursor to the start of the buffer after each group's matmul, so groups do not accumulate.
:::
:::note
Callers reaching the kernel through `arm_convolve_wrapper_s8()` must size the buffer with `arm_convolve_wrapper_s8_get_buffer_size()` instead, which covers every kernel the wrapper may dispatch to. This function is for callers that invoke `arm_convolve_1x1_out_s8()` directly.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, KH, KW, C_IN] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | For valid (non-negative, in-range) filter dimensions, the buffer size in bytes: round_up_4(KH * KW * C_IN) on builds with the MVE extension (ARM_MATH_MVEI), 0 otherwise, since `arm_convolve_1x1_out_s8()` only exists on MVE builds. Returns -1 if any of filter_dims->w, filter_dims->h or filter_dims->c is negative or out of int32_t range, or if the rounded-up product exceeds INT32_MAX. The validation runs on every build target, not just the MVE leg, so the contract does not vary by target. |
Source: `Include/arm_nnfunctions.h:1554`
## arm_convolve_1_x_n_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_convolve_1_x_n_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
1xn convolution for s4 weights
- Supported framework : TensorFlow Lite Micro
- The following constrains on the arguments apply
1. stride.w * input_dims->c is a multiple of 4
2. Explicit constraints(since it is for 1xN convolution) -## input_dims->h equals 1 -## output_dims->h equals 1 -## filter_dims->h equals 1
:::note[Todo]
Remove constraint on output_dims->w to make the function generic.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_1_x_n_s4_get_buffer_size will return the buffer_size if required The caller is expected to clear the buffer, if applicable, for security reasons. |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 as packed int4 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:1592`
## arm_convolve_1_x_n_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_1_x_n_s8_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required additional buffer size for 1xn convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or conv_params->stride.w is not positive. On builds with the MVE extension (ARM_MATH_MVEI) that is the staging size of `arm_convolve_1_x_n_s8()`, at least filter W * C_IN bytes, or -1 if it would not fit in an int32_t; other builds return `arm_convolve_s8_get_buffer_size()`. |
Source: `Include/arm_nnfunctions.h:1621`
## arm_convolve_1_x_n_s4_get_buffer_size
`function` · `c`
```c
int32_t arm_convolve_1_x_n_s4_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get the required additional buffer size for 1xn convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or conv_params->stride.w is not positive. It also returns -1 if the required size would not fit in an int32_t; on a Helium build the route whose padding lines up with the stride needs no buffer and returns 0 without computing one. |
Source: `Include/arm_nnfunctions.h:1643`
## arm_depthwise_conv_wrapper_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_wrapper_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Wrapper function to pick the right optimized s8 depthwise convolution function.
- Supported framework: TensorFlow Lite
- Picks one of the the following functions
1. `arm_depthwise_conv_s8()`
2. `arm_depthwise_conv_3x3_s8()` - Cortex-M CPUs with DSP extension only
3. `arm_depthwise_conv_s8_opt()`
- Check details of `arm_depthwise_conv_s8_opt()` for potential data that can be accessed outside of the boundary.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s8_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s8_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer, if applicable, for security reasons. |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, supplied by the caller. The selected kernel only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter, bias and dw_conv_params->input_offset are unchanged - see `arm_depthwise_convolve_weight_sum()` for the layout and the full reuse rules. Whether the buffer is consumed at all depends on the route this wrapper takes. It is forwarded to `arm_depthwise_conv_s8_opt()`, which reads it under MVE, only when dw_conv_params->ch_mult == 1, input_dims->n == 1, and either both dilations are 1 or the layer is 1D and dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, dw_conv_params->padding.h == 0, dilation.h == 1 and dilation.w >= 1. Such a dilated 1D layer therefore reads the sums too. Outside those cases the wrapper calls `arm_depthwise_conv_s8()`, which has no such parameter and ignores the context entirely - which is why several in-tree tests legitimately pass sums built by `arm_convolve_weight_sum()`, or none at all, on those routes (a 2D-dilated layer, for example). On MVE with input_dims->c == 1 and an output channel count above CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (8 on armclang, 1 otherwise), the layer is instead converted to a regular convolution, and conv-style sums from `arm_convolve_weight_sum()` are what that route wants. Where the sums are actually read, fill the buffer with `arm_depthwise_convolve_weight_sum()`, passing dw_conv_params->input_offset as lhs_offset and the same bias given here, so that entry j holds input_offset * sum(weights of channel j) + bias[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. On the `arm_depthwise_conv_s8_opt()` route, a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on builds where the buffer is actually read (ARM_MATH_DSP and ARM_MATH_MVEI both defined); on other builds the parameter is unread and NULL is accepted. On MVE with input_dims->c == 1 and an output channel count above CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (8 on armclang, 1 otherwise), this wrapper instead diverts to `arm_convolve_wrapper_s8()`. That diversion exists only on MVE, and every kernel it can dispatch to diagnoses a NULL buf with ARM_CMSIS_NN_ARG_ERROR, so that route is covered too. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions; see weight_sum_ctx for which dilated layers take the `arm_depthwise_conv_s8_opt()` route. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` on successful completion, or `ARM_CMSIS_NN_ARG_ERROR` on the `arm_depthwise_conv_s8_opt()` route if ctx->buf is NULL when a scratch buffer is required, or if weight_sum_ctx->buf is NULL on builds where it is read (ARM_MATH_DSP and ARM_MATH_MVEI both defined), or if ctx->size is non-zero and below `arm_depthwise_conv_s8_opt_get_buffer_size()` for a layer its channel path runs, or if that sizer returns -1 (a negative dimension or a byte count it cannot represent), or on the MVE `arm_convolve_wrapper_s8()` diversion route if weight_sum_ctx->buf is NULL. |
Source: `Include/arm_nnfunctions.h:1731`
## arm_depthwise_conv_wrapper_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_wrapper_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Wrapper function to pick the right optimized s4 depthwise convolution function.
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s4_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s4_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer ,if applicable, for security reasons. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43]. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` - Successful completion. |
Source: `Include/arm_nnfunctions.h:1779`
## arm_depthwise_conv_wrapper_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get size of additional buffer required by `arm_depthwise_conv_wrapper_s8()`.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:1818`
## arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get size of additional buffer required by `arm_depthwise_conv_wrapper_s8()` for processors with DSP extension.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s8_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:1834`
## arm_depthwise_conv_wrapper_s8_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size_mve(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get size of additional buffer required by `arm_depthwise_conv_wrapper_s8()` for Arm(R) Helium Architecture case.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s8_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:1850`
## arm_depthwise_conv_wrapper_s4_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get size of additional buffer required by `arm_depthwise_conv_wrapper_s4()`.
This sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as `arm_depthwise_conv_s8_opt_get_buffer_size()`, so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:1878`
## arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get size of additional buffer required by `arm_depthwise_conv_wrapper_s4()` for processors with DSP extension.
This sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as `arm_depthwise_conv_s8_opt_get_buffer_size()`, so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s4_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked. This variant forwards to the top-level dispatcher, so it follows the build's leg; both legs inspect input_dims->c.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:1895`
## arm_depthwise_conv_wrapper_s4_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size_mve(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get size of additional buffer required by `arm_depthwise_conv_wrapper_s4()` for Arm(R) Helium Architecture case.
This sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as `arm_depthwise_conv_s8_opt_get_buffer_size()`, so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s4_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked. The Helium leg sizes its buffer from a fixed channel block rather than from input_dims->c, but it checks that dimension anyway so that this variant answers a negative channel count with the same -1 the dispatcher returns (issue #318).
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:1913`
## arm_depthwise_conv_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Basic s8 depthwise convolution function that doesn't have any constraints on the input dimensions.
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s8_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s8()`, size the context with `arm_depthwise_conv_wrapper_s8_get_buffer_size()` instead, because another route through that wrapper does require a buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:1947`
## arm_depthwise_conv_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input,
const cmsis_nn_dims *filter_dims,
const int8_t *kernel,
const cmsis_nn_dims *bias_dims,
const int32_t *bias,
const cmsis_nn_dims *output_dims,
int8_t *output
)
```
Basic s4 depthwise convolution function that doesn't have any constraints on the input dimensions.
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s4_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s4()`, size the context with `arm_depthwise_conv_wrapper_s4_get_buffer_size()` instead, because another route through that wrapper does require a buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
| input | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| kernel | const int8_t * | in | Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43]. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output | int8_t * | in, out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:1989`
## arm_depthwise_conv_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
Basic s16 depthwise convolution function that doesn't have any constraints on the input dimensions.
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s16_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s16()`, size the context with `arm_depthwise_conv_wrapper_s16_get_buffer_size()` instead, because another route through that wrapper does require a buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) conv_params->input_offset : Not used conv_params->output_offset : Not used |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:2029`
## arm_depthwise_conv_wrapper_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_wrapper_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
Wrapper function to pick the right optimized s16 depthwise convolution function.
- Supported framework: TensorFlow Lite
- Picks one of the the following functions
1. `arm_depthwise_conv_s16()`
2. `arm_depthwise_conv_fast_s16()` - Cortex-M CPUs with DSP extension only
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s16_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s16_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer, if applicable, for security reasons. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions. When ch_mult == 1 and filter_dims->w * filter_dims->h < 512, `arm_depthwise_conv_fast_s16()` is used for an undilated layer and for a 1D layer dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, dw_conv_params->padding.h == 0, dilation.h == 1 and dilation.w >= 1. Other layers use `arm_depthwise_conv_s16()`. Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->output_offset : Not used |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` - Successful completion. |
Source: `Include/arm_nnfunctions.h:2082`
## arm_depthwise_conv_wrapper_s16_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get size of additional buffer required by `arm_depthwise_conv_wrapper_s16()`.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:2119`
## arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get size of additional buffer required by `arm_depthwise_conv_wrapper_s16()` for processors with DSP extension.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s16_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:2135`
## arm_depthwise_conv_wrapper_s16_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size_mve(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Get size of additional buffer required by `arm_depthwise_conv_wrapper_s16()` for Arm(R) Helium Architecture case.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s16_get_buffer_size()`.
:::
:::note
An out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Source: `Include/arm_nnfunctions.h:2152`
## arm_depthwise_conv_fast_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_fast_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
Optimized s16 depthwise convolution function with constraint that in_channel equals out_channel.
`ARM_CMSIS_NN_SUCCESS` - Successful operation
- Supported framework: TensorFlow Lite
- The following constraints on the arguments apply
1. ch_mult == 1: the number of input channels equals the number of output channels
2. filter_dims->w * filter_dims->h < MAX_COL_COUNT (512)
3. dw_conv_params->dilation.h == 1 and dw_conv_params->dilation.w >= 1
- Recommended when number of channels is 4 or greater.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. `arm_depthwise_conv_fast_s16_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation.w is honoured; dw_conv_params->dilation.h must be 1. dw_conv_params->input_offset : Not used dw_conv_params->output_offset : Not used |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - ctx-buff == NULL and `arm_depthwise_conv_fast_s16_get_buffer_size()` != 0 or input channel != output channel or filter_dims->w * filter_dims->h >= MAX_COL_COUNT (512) or dw_conv_params->dilation.h != 1 or dw_conv_params->dilation.w < 1 |
Source: `Include/arm_nnfunctions.h:2200`
## arm_depthwise_conv_fast_s16_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_fast_s16_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for optimized s16 depthwise convolution function with constraint that in_channel equals out_channel.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:2225`
## arm_depthwise_conv_3x3_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_3x3_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below).
- Supported framework : TensorFlow Lite Micro
- The following constrains on the arguments apply
1. Number of input channel equals number of output channels
2. Filter height and width equals 3
3. Padding along x is either 0 or 1.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - Unsupported dimension of tensors
- Unsupported pad size along the x axis `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:2261`
## arm_depthwise_conv_s8_opt
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_s8_opt(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Optimized s8 depthwise convolution function with constraint that in_channel equals out_channel.
:::note
The second argument, weight_sum_ctx, has no counterpart on `arm_depthwise_conv_s8()`, so it is described here rather than by reference. It carries per-channel weight sums that the caller supplies: this function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as the weights, the bias and dw_conv_params->input_offset are unchanged - see `arm_depthwise_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_depthwise_convolve_weight_sum()`, which walks the channel-interleaved depthwise weight layout; `arm_convolve_weight_sum()` sums a different set of weights and is not a substitute here. That helper returns `ARM_CMSIS_NN_NO_IMPL_ERROR` on non-MVE builds, which is not a failure. Size the buffer with `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. Clear the buffer afterwards if applicable for security reasons. Pass a valid context on every build. On builds where the buffer is actually read (ARM_MATH_DSP and ARM_MATH_MVEI both defined), a NULL buf is diagnosed and this function returns `ARM_CMSIS_NN_ARG_ERROR`, matching `arm_convolve_s8()`. On other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: it still produces wrong output while returning `ARM_CMSIS_NN_SUCCESS`, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions.
:::
:::note
ctx->size is optional: a caller that leaves it at zero opts out of the size check, as TFLM does. On the channel path, a non-zero ctx->size below `arm_depthwise_conv_s8_opt_get_buffer_size()` is rejected before any write. A layer the planar path takes needs only its plane, so it can succeed with less.
:::
:::note
MVE channel tail loads and stores are predicated, so channel-indexed arrays are not accessed beyond the number of channels.
:::
- Supported framework: TensorFlow Lite
- The following constrains on the arguments apply
1. Number of input channel equals number of output channels or ch_mult equals 1
- Reccomended when number of channels is 4 or greater.
- On builds with ARM_MATH_DSP and ARM_MATH_MVEI, layers that `arm_depthwise_conv_s8_opt_planar_supported()` accepts run the planar path, with the same result as `arm_depthwise_conv_s8_opt_planar()`, unless ctx->size cannot hold its plane; every other layer runs the channel path of `arm_depthwise_conv_s8_opt_channelwise()`. Callers that choose the path ahead of time can call either one directly.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. `arm_depthwise_conv_s8_opt_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, supplied by the caller and only read by this function. See the note below for how to size, fill and reuse the buffer and for when a NULL buf is diagnosed. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation.w is honoured; dw_conv_params->dilation.h must be 1. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - input channel != output channel or ch_mult != 1, or dw_conv_params->dilation.h != 1 or dw_conv_params->dilation.w < 1, or ctx->buf is NULL when a scratch buffer is required, or ctx->size is non-zero and below `arm_depthwise_conv_s8_opt_get_buffer_size()` for a layer the channel path runs, or that sizer returns -1 (a negative dimension or a byte count it cannot represent) on the channel path, or weight_sum_ctx->buf is NULL on builds where it is read (ARM_MATH_DSP and ARM_MATH_MVEI both defined) `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:2347`
## arm_depthwise_conv_s8_opt_planar_supported
`function` · `c`
```c
int32_t arm_depthwise_conv_s8_opt_planar_supported(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Whether `arm_depthwise_conv_s8_opt()` runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels.
- The rule is plain C and evaluates the same on every build, so a code generator can apply it ahead of time. The planar path itself exists only on builds with ARM_MATH_DSP and ARM_MATH_MVEI.
- It depends only on the shapes and dw_conv_params: batch 1, C_IN equal to C_OUT, positive dimensions, stride 1, ch_mult 1, dilation.h 1, dilation.w at least 1 and at most 128 / C (integer division), at most 32 channels, the widths the path is faster for, and a plane that fits the scratch. This function is the reference for the rule; a mirror should be checked against it.
- The width thresholds follow measured speed and may be retuned in a later release. A caller that calls `arm_depthwise_conv_s8_opt_planar()` directly must handle `ARM_CMSIS_NN_NO_IMPL_ERROR`, for example by calling `arm_depthwise_conv_s8_opt_channelwise()`.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 1 when the planar path takes the layer with a scratch of `arm_depthwise_conv_s8_opt_get_buffer_size_mve()` bytes, 0 otherwise. |
Source: `Include/arm_nnfunctions.h:2384`
## arm_depthwise_conv_s8_opt_planar
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_s8_opt_planar(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
The planar path of `arm_depthwise_conv_s8_opt()` on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane.
- The output is identical to `arm_depthwise_conv_s8_opt()` and `arm_depthwise_conv_s8()`. The bias is read through the weight sums; bias_dims and bias_data are unused.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context with the `arm_depthwise_conv_s8_opt_get_buffer_size()` scratch |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()` |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()` |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_depthwise_conv_s8_opt()`, except its channel-path ctx->size check: a ctx->size too small for the plane returns ARM_CMSIS_NN_NO_IMPL_ERROR instead `ARM_CMSIS_NN_NO_IMPL_ERROR` - `arm_depthwise_conv_s8_opt_planar_supported()` rejects the layer, ctx->size cannot hold its plane, or the build lacks ARM_MATH_DSP or ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:2420`
## arm_depthwise_conv_s8_opt_channelwise
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_s8_opt_channelwise(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
The channel-vectorized path of `arm_depthwise_conv_s8_opt()` on its own, without the planar attempt.
- The output is identical to `arm_depthwise_conv_s8_opt()` and `arm_depthwise_conv_s8()` for every layer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context with the `arm_depthwise_conv_s8_opt_get_buffer_size()` scratch |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()` |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()` |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_depthwise_conv_s8_opt()` `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:2457`
## arm_depthwise_conv_s8_opt_3x3
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_s8_opt_3x3(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of `arm_depthwise_conv_s8_opt()`, for layers in its gate. It computes three output rows per weight load.
- The output is identical to `arm_depthwise_conv_s8_opt()` and `arm_depthwise_conv_s8()`. The bias is read through the weight sums, which `arm_depthwise_convolve_weight_sum()` fills as for `arm_depthwise_conv_s8_opt()`; bias_dims and bias_data are unused.
- Gate: filter 3x3, dilation 1, N 1, C_IN equal to C_OUT with 16 <= C <= 2048 and C % 4 == 0, stride 1 or 2 and padding 0 or 1 in each dimension, input W >= 3 and H >= 1, output H >= 3 and W x H >= 16, every dimension at most 4096 and each tensor at most INT32_MAX elements, and the centre of the last output column's window inside the input: (output W - 1) x stride.w - padding.w + 1 < input W. Input rows above or below the input count as padding, as in `arm_depthwise_conv_s8()`.
- Buffers: ctx and weight_sum_ctx and their buf are non-NULL, and ctx->size is at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` (3008 + input W x C + 16 bytes). ctx->buf needs no alignment.
- It is a direct entry: `arm_depthwise_conv_s8_opt()` does not call it. A caller that selects the kernel per layer ahead of time calls `arm_depthwise_conv_s8_opt_3x3_c64_s1()` for C 64 with stride.h 1, this function for the rest of the gate, and `arm_depthwise_conv_s8_opt()` for every other layer, or on `ARM_CMSIS_NN_NO_IMPL_ERROR`. All three take the same arguments and the same weight sums; a ctx that serves all three holds the larger of `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` and `arm_depthwise_conv_s8_opt_get_buffer_size()`. On ARM_MATH_MVEI builds the second is the larger for a 3x3 filter when input W x C <= 1440.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context with at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` bytes of scratch |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()` |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()` |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer or a buffer is outside the gate below, or the build lacks ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:2513`
## arm_depthwise_conv_s8_opt_3x3_c64_s1
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_s8_opt_3x3_c64_s1(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
`arm_depthwise_conv_s8_opt_3x3()` specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps. It is the faster entry for those layers.
- The output is identical to `arm_depthwise_conv_s8_opt_3x3()`, `arm_depthwise_conv_s8_opt()` and `arm_depthwise_conv_s8()`. The bias is read through the weight sums; bias_dims and bias_data are unused.
- The two entries share only their gate, parameter packing and the code for the output_y % 3 remainder rows, so a build with -ffunction-sections and section garbage collection keeps only the code of the entries it calls.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context with at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` bytes of scratch |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()` |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()` |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_NO_IMPL_ERROR` - C_IN is not 64, dw_conv_params->stride.h is not 1, the layer or a buffer is outside the gate of `arm_depthwise_conv_s8_opt_3x3()`, or the build lacks ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:2558`
## arm_depthwise_conv_s8_opt_3x3_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_s8_opt_3x3_get_buffer_size(const cmsis_nn_dims *input_dims)
```
Get the scratch size in bytes of `arm_depthwise_conv_s8_opt_3x3()` and `arm_depthwise_conv_s8_opt_3x3_c64_s1()`.
- The size is the minimum ctx->size both entries accept. It depends only on the input width and channel count, since the filter is always 3x3.
- The function is plain C and returns the same size on every build, including builds without ARM_MATH_MVEI where the entries return `ARM_CMSIS_NN_NO_IMPL_ERROR`, so a code generator can size the scratch ahead of time. A non-negative size is not a statement that the layer is in the gate of the entries.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN]. Only W and C_IN are read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 3008 + W x C_IN + 16 bytes, or -1 if W or C_IN is negative or the size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:2585`
## arm_depthwise_conv_s4_opt
`function` · `c`
```c
arm_cmsis_nn_status arm_depthwise_conv_s4_opt(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Optimized s4 depthwise convolution function with constraint that in_channel equals out_channel.
:::note
MVE channel tail loads and stores are predicated, so channel-indexed arrays are not accessed beyond the number of channels.
:::
- Supported framework: TensorFlow Lite
- The following constrains on the arguments apply
1. Number of input channel equals number of output channels or ch_mult equals 1
- Reccomended when number of channels is 4 or greater.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer required by the function. `arm_depthwise_conv_s4_opt_get_buffer_size()` will return the buffer_size. A NULL ctx->buf is diagnosed with `ARM_CMSIS_NN_ARG_ERROR`. The caller is expected to clear the buffer, if applicable, for security reasons. |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43]. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - input channel != output channel or ch_mult != 1 `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:2626`
## arm_depthwise_conv_s8_opt_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_s8_opt_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:2651`
## arm_depthwise_conv_s4_opt_get_buffer_size
`function` · `c`
```c
int32_t arm_depthwise_conv_s4_opt_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for optimized s4 depthwise convolution function with constraint that in_channel equals out_channel.
The dimensions are not checked here: the query routes straight to the s8 _mve/_dsp leg and relies on the range checks inside that leg. Both legs apply the same check as `arm_depthwise_conv_s8_opt_get_buffer_size()`, so the answer for an out-of-range shape is the same on every build target.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if input_dims->c or a filter dimension it reads is negative, or the required size would not fit in an int32_t. |
Source: `Include/arm_nnfunctions.h:2667`
## arm_fully_connected_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Basic s4 Fully Connected function.
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_fully_connected_s4_get_buffer_size(). Do not size this context with `arm_fully_connected_s8_get_buffer_size()`: that sizes the kernel-sum buffer of a different kernel and does not describe this argument. |
| fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_tensor_quant_params * | in | Per-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor. |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43]. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:2717`
## arm_fully_connected_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Basic s8 Fully Connected function.
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Per-output-channel kernel sums, supplied by the caller - not scratch memory that this function fills in. The library never populates ctx->buf for this function, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds this function adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. Sized by `arm_fully_connected_s8_get_buffer_size()`: filter_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call. |
| fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_tensor_quant_params * | in | Per-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor. |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:2789`
## arm_fully_connected_per_channel_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_per_channel_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
Basic s8 Fully Connected function using per channel quantization.
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Per-output-channel kernel sums, supplied by the caller - not scratch memory that this function fills in. The library never populates ctx->buf for this function, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds this function adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. There is no per-channel sizer; `arm_fully_connected_s8_get_buffer_size()` returns the same quantity this function needs, filter_dims->c * sizeof(int32_t) where the sums are used and 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call. |
| fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:2862`
## arm_fully_connected_wrapper_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_wrapper_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 Fully Connected layer wrapper function
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Per-output-channel kernel sums, supplied by the caller - not scratch memory that this wrapper fills in. The library never populates ctx->buf here, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. The context is passed straight through to `arm_fully_connected_per_channel_s8()` or `arm_fully_connected_s8()` depending on quant_params->is_per_channel, and both read it the same way. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. Neither selected kernel writes the buffer, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds the selected kernel adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. There is no wrapper sizer; `arm_fully_connected_s8_get_buffer_size()` returns the quantity both routes need, filter_dims->c * sizeof(int32_t) where the sums are used and 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call. |
| fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127] |
| quant_params | const cmsis_nn_quant_params * | in | Per-channel or per-tensor quantization info. Check struct defintion for details. It contains the multiplier and shift value(s) to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
| bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:2937`
## arm_vector_sum_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_vector_sum_s8(
int32_t *vector_sum_buf,
const int32_t vector_cols,
const int32_t vector_rows,
const int8_t *vector_data,
const int32_t lhs_offset,
const int32_t rhs_offset,
const int32_t *bias_data
)
```
Calculate the sum of each row in vector_data, multiply by lhs_offset and optionally add s32 bias_data.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| vector_sum_buf | int32_t * | in, out | Buffer for vector sums |
| vector_cols | const int32_t | in | Number of vector columns |
| vector_rows | const int32_t | in | Number of vector rows |
| vector_data | const int8_t * | in | Vector of weigths data |
| lhs_offset | const int32_t | in | Constant multiplied with each sum |
| rhs_offset | const int32_t | in | Constant added to each vector element before sum |
| bias_data | const int32_t * | in | Vector of bias data, added to each sum. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:2961`
## arm_vector_sum_s8_s64
`function` · `c`
```c
arm_cmsis_nn_status arm_vector_sum_s8_s64(
int64_t *vector_sum_buf,
const int32_t vector_cols,
const int32_t vector_rows,
const int8_t *vector_data,
const int32_t lhs_offset,
const int64_t *bias_data
)
```
Calculate the sum of each row in vector_data, multiply by lhs_offset and optionally add s64 bias_data.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| vector_sum_buf | int64_t * | in, out | Buffer for vector sums |
| vector_cols | const int32_t | in | Number of vector columns |
| vector_rows | const int32_t | in | Number of vector rows |
| vector_data | const int8_t * | in | Vector of weigths data |
| lhs_offset | const int32_t | in | Constant multiplied with each sum |
| bias_data | const int64_t * | in | Vector of bias data, added to each sum. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:2980`
## arm_fully_connected_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_fully_connected_s8_get_buffer_size(const cmsis_nn_dims *filter_dims)
```
Get size of additional buffer required by `arm_fully_connected_s8()`. See also arm_vector_sum_s8, which is required if buffer size is > 0.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | dimension of filter |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:2998`
## arm_fully_connected_s8_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_fully_connected_s8_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)
```
Get size of additional buffer required by `arm_fully_connected_s8()` for processors with DSP extension.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_s8_get_buffer_size()`.
:::
:::note
Also validates dims like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | dimension of filter |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:3009`
## arm_fully_connected_s8_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_fully_connected_s8_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)
```
Get size of additional buffer required by `arm_fully_connected_s8()` for Arm(R) Helium Architecture case.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_s8_get_buffer_size()`.
:::
:::note
Also validates dims like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | dimension of filter |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:3020`
## arm_fully_connected_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
Basic s16 Fully Connected function.
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused. This function currently ignores the context entirely on every build - it neither reads nor writes ctx->buf - and `arm_fully_connected_s16_get_buffer_size()` returns 0 accordingly, so { NULL, 0 } is accepted. Unlike the s8 variants, no precomputed kernel sums are required here. None of this is a guarantee about future versions. |
| fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. fc_params->input_offset : 0 fc_params->filter_offset : 0 fc_params->output_offset : 0 |
| quant_params | const cmsis_nn_per_tensor_quant_params * | in | Per-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor. |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
| bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:3057`
## arm_fully_connected_per_channel_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_per_channel_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *kernel,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
Basic s16 Fully Connected function using per channel quantization.
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Scratch buffer that this function writes before it reads, on every build. It is filled here with one reduced int32 multiplier per output channel derived from quant_params->multiplier, so the caller supplies the storage only and the incoming contents are never used. Unlike the s8 variants, no precomputed kernel sums are expected, and clearing the buffer is harmless. Required on every build, not only under MVE: ctx or ctx->buf being NULL is reported as ARM_CMSIS_NN_ARG_ERROR, as is a non-zero ctx->size smaller than the requirement. A ctx->size of 0 is treated as undeclared and is not checked. Sized by `arm_fully_connected_per_channel_s16_get_buffer_size()`: filter_dims->c * sizeof(int32_t), which equals the output_dims->c entries written. The caller is expected to clear the buffer afterwards, if applicable, for security reasons. |
| fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : 0 fc_params->filter_offset : 0 Range of fc_params->output_offset : 0 |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
| kernel | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
| bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:3114`
## arm_fully_connected_wrapper_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_fully_connected_wrapper_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
s16 Fully Connected layer wrapper function
- Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Scratch buffer, whose use depends on the route taken. Unlike the s8 wrapper, no precomputed kernel sums are expected on either route, and clearing the buffer is harmless. When quant_params->is_per_channel is set, the context is passed to `arm_fully_connected_per_channel_s16()`, which writes it before reading it, on every build: it is filled there with one reduced int32 multiplier per output channel, so the caller supplies the storage only. On that route ctx or ctx->buf being NULL is reported as ARM_CMSIS_NN_ARG_ERROR, as is a non-zero ctx->size smaller than filter_dims->c * sizeof(int32_t); a ctx->size of 0 is treated as undeclared and is not checked. Size it with `arm_fully_connected_per_channel_s16_get_buffer_size()`. Otherwise the context goes to `arm_fully_connected_s16()`, which currently ignores it entirely, so { NULL, 0 } is accepted on that route. A caller that does not know the route in advance should size for the per-channel case, since `arm_fully_connected_s16_get_buffer_size()` returns 0. None of this is a guarantee about future versions. The caller is expected to clear the buffer afterwards, if applicable, for security reasons. |
| fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : 0 fc_params->filter_offset : 0 Range of fc_params->output_offset : 0 |
| quant_params | const cmsis_nn_quant_params * | in | Per-channel or per-tensor quantization info. Check struct defintion for details. It contains the multiplier and shift value(s) to be applied to each output channel |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
| filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
| bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:3174`
## arm_fully_connected_s16_get_buffer_size
`function` · `c`
```c
int32_t arm_fully_connected_s16_get_buffer_size(const cmsis_nn_dims *filter_dims)
```
Get size of additional buffer required by `arm_fully_connected_s16()`.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | dimension of filter |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes |
Source: `Include/arm_nnfunctions.h:3192`
## arm_fully_connected_s16_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_fully_connected_s16_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)
```
Get size of additional buffer required by `arm_fully_connected_s16()` for processors with DSP extension.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_s16_get_buffer_size()`.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | dimension of filter |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes |
Source: `Include/arm_nnfunctions.h:3202`
## arm_fully_connected_s16_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_fully_connected_s16_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)
```
Get size of additional buffer required by `arm_fully_connected_s16()` for Arm(R) Helium Architecture case.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_s16_get_buffer_size()`.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | dimension of filter |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes |
Source: `Include/arm_nnfunctions.h:3212`
## arm_fully_connected_per_channel_s16_get_buffer_size
`function` · `c`
```c
int32_t arm_fully_connected_per_channel_s16_get_buffer_size(const cmsis_nn_dims *filter_dims)
```
Get size of additional buffer required by `arm_fully_connected_per_channel_s16()`.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | dimension of filter |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:3223`
## arm_fully_connected_per_channel_s16_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_fully_connected_per_channel_s16_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)
```
Get size of additional buffer required by `arm_fully_connected_per_channel_s16()` for processors with DSP extension.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_per_channel_s16_get_buffer_size()`.
:::
:::note
Also validates dims like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | dimension of filter |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:3235`
## arm_fully_connected_per_channel_s16_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_fully_connected_per_channel_s16_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)
```
Get size of additional buffer required by `arm_fully_connected_per_channel_s16()` for Arm(R) Helium Architecture case.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_per_channel_s16_get_buffer_size()`.
:::
:::note
Also validates dims like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| filter_dims | const cmsis_nn_dims * | in | dimension of filter |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:3247`
## arm_add_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_add_s8(
const int8_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int8_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)
```
s8 elementwise add of two tensors with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input1_data | const int8_t * | in | pointer to input tensor 1 |
| input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
| input2_data | const int8_t * | in | pointer to input tensor 2 |
| input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
| input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input1_mult | const int32_t | in | multiplier for input 1 |
| input1_shift | const int32_t | in | shift for input 1 |
| input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input2_mult | const int32_t | in | multiplier for input 2 |
| input2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output_data | int8_t * | out | pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
Source: `Include/arm_nnfunctions.h:3285`
## arm_add_scalar_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_add_scalar_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s8 elementwise add of scalar and vector
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int8_t * | in | pointer to input scalar |
| input_2_vect | const int8_t * | in | pointer to input vector |
| input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int8_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3329`
## arm_elementwise_add_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_add_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s8 elementwise add of two vectors
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int8_t * | in | pointer to input vector 1 |
| input_2_vect | const int8_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | input left shift. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int8_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3369`
## arm_abs_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_abs_s8(
const int8_t *input,
const int32_t input_offset,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const bool needs_rescale,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s8 elementwise absolute value
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | pointer to input vector |
| input_offset | const int32_t | in | input offset |
| output | int8_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| needs_rescale | const bool | in | indicates if output requantization is needed |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3400`
## arm_sqrt_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_sqrt_s8(
const int8_t *input,
const cmsis_nn_dims *input_dims,
int8_t *output,
const int8_t *sqrt_lut
)
```
s8 elementwise square root
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | pointer to input vector |
| input_dims | const cmsis_nn_dims * | in | pointer to input tensor dimensions |
| output | int8_t * | out | pointer to output vector |
| sqrt_lut | const int8_t * | in | pointer to 256-entry lookup table |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3420`
## arm_sqrt_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_sqrt_s16(
const int16_t *input,
const cmsis_nn_dims *input_dims,
int16_t *output,
const int16_t *sqrt_lut
)
```
s16 elementwise square root using piecewise LUT with linear interpolation
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | pointer to input vector |
| input_dims | const cmsis_nn_dims * | in | pointer to input tensor dimensions |
| output | int16_t * | out | pointer to output vector |
| sqrt_lut | const int16_t * | in | pointer to 513-entry lookup table (int16_t) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3431`
## arm_sqrt_s16_tablefree
`function` · `c`
```c
arm_cmsis_nn_status arm_sqrt_s16_tablefree(
const int16_t *input,
const cmsis_nn_dims *input_dims,
int16_t *output,
const float scale
)
```
s16 elementwise square root without a lookup table
Approximates output[i] = trunc(sqrt(input[i] * scale)) saturated to 32767, which is LiteRT's int16 SQRT (dequantize in float32, sqrtf, divide by the output scale, truncate, clamp) for zero points 0, to within 1 LSB of LiteRT at every non-negative input for input scales 1e-7 to 1e-1 and output scales from 0.01x to 10x the full-range scale, saturating ones included. Inputs at or below 0 produce
1. Needs no table; the int16 API does not depend on ARM_NN_ENABLE_F32/F16, and on targets without a floating-point unit the plain C path uses fmaf from the C library.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | pointer to input vector |
| input_dims | const cmsis_nn_dims * | in | pointer to input tensor dimensions |
| output | int16_t * | out | pointer to output vector |
| scale | const float | in | input_scale / (output_scale * output_scale) as float32: take the float32-rounded tensor scales, evaluate in float64 and round once to float32. Must be finite and greater than 0. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3454`
## arm_abs_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_abs_s16(
const int16_t *input,
const int32_t input_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const bool needs_rescale,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise absolute value
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | pointer to input vector |
| input_offset | const int32_t | in | input offset |
| output | int16_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| needs_rescale | const bool | in | indicates if output requantization is needed |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3470`
## arm_rsqrt_s16_per_op
`function` · `c`
```c
arm_cmsis_nn_status arm_rsqrt_s16_per_op(
const int16_t *input,
const int32_t input_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size,
const int16_t *lut
)
```
INT16 reciprocal square root using a per-operator LUT.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input buffer. |
| input_offset | const int32_t | in | Input tensor zero offset. The kernel evaluates each element as `input - input_offset` before the LUT lookup. |
| output | int16_t * | out | Pointer to the output buffer. |
| out_offset | const int32_t | in | Output tensor zero offset. |
| out_activation_min | const int32_t | in | Minimum output clamp. |
| out_activation_max | const int32_t | in | Maximum output clamp. |
| block_size | const int32_t | in | Number of elements. |
| lut | const int16_t * | in | Pointer to a 513-entry INT16 LUT in output domain. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions.h:3496`
## arm_rsqrt_s16_universal
`function` · `c`
```c
arm_cmsis_nn_status arm_rsqrt_s16_universal(
const int16_t *input,
const int32_t input_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const bool needs_rescale,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size,
const int32_t *lut
)
```
INT16 reciprocal square root using a shared universal LUT.
In universal mode all RSQRT operators share a single LUT that captures the base 1/sqrt(x) shape, and operator-specific quantization is applied afterward via `out_mult` / `out_shift`. Because this two-step process introduces extra rounding stages, the output may differ from the per-op variant (`arm_rsqrt_s16_per_op`) by up to ±3 LSB per element. This is expected and acceptable for deployment.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input buffer. |
| input_offset | const int32_t | in | Input tensor zero offset. The kernel evaluates each element as `input - input_offset` before the LUT lookup. |
| output | int16_t * | out | Pointer to the output buffer. |
| out_offset | const int32_t | in | Output tensor zero offset. |
| out_mult | const int32_t | in | Output requantization multiplier. |
| out_shift | const int32_t | in | Output requantization shift. |
| needs_rescale | const bool | in | Whether requantization is required. |
| out_activation_min | const int32_t | in | Minimum output clamp. |
| out_activation_max | const int32_t | in | Maximum output clamp. |
| block_size | const int32_t | in | Number of elements. |
| lut | const int32_t * | in | Pointer to a 513-entry INT32 shared LUT in Q30 domain. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
Source: `Include/arm_nnfunctions.h:3530`
## arm_sub_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_sub_s8(
const int8_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int8_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)
```
s8 elementwise subtraction of two tensors with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input1_data | const int8_t * | in | pointer to input tensor 1 |
| input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
| input2_data | const int8_t * | in | pointer to input tensor 2 |
| input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
| input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input1_mult | const int32_t | in | multiplier for input 1 |
| input1_shift | const int32_t | in | shift for input 1 |
| input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input2_mult | const int32_t | in | multiplier for input 2 |
| input2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output_data | int8_t * | out | pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
Source: `Include/arm_nnfunctions.h:3571`
## arm_sub_scalar_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_sub_scalar_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s8 elementwise subtract of scalar and vector (scalar - vector)
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int8_t * | in | pointer to input scalar |
| input_2_vect | const int8_t * | in | pointer to input vector |
| input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int8_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3615`
## arm_elementwise_sub_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_sub_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s8 elementwise subtract of two vectors
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int8_t * | in | pointer to input vector 1 |
| input_2_vect | const int8_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | input left shift. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int8_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3655`
## arm_add_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_add_s16(
const int16_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int16_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)
```
s16 elementwise add of two tensors with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input1_data | const int16_t * | in | pointer to input tensor 1 |
| input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
| input2_data | const int16_t * | in | pointer to input tensor 2 |
| input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
| input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input1_mult | const int32_t | in | multiplier for input 1 |
| input1_shift | const int32_t | in | shift for input 1 |
| input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input2_mult | const int32_t | in | multiplier for input 2 |
| input2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output_data | int16_t * | out | pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
Source: `Include/arm_nnfunctions.h:3702`
## arm_add_scalar_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_add_scalar_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise add of scalar and vector
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input scalar |
| input_2_vect | const int16_t * | in | pointer to input vector |
| input_1_offset | const int32_t | in | offset for input 1. Not used. |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Not used. |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int16_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Not used. |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3747`
## arm_elementwise_add_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_add_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise add of two vectors
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1. Not used. |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Not used. |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | input left shift. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int16_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Not used. |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3789`
## arm_sub_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_sub_s16(
const int16_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int16_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)
```
s16 elementwise subtraction of two tensors with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input1_data | const int16_t * | in | pointer to input tensor 1 |
| input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
| input2_data | const int16_t * | in | pointer to input tensor 2 |
| input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
| input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input1_mult | const int32_t | in | multiplier for input 1 |
| input1_shift | const int32_t | in | shift for input 1 |
| input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input2_mult | const int32_t | in | multiplier for input 2 |
| input2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output_data | int16_t * | out | pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
Source: `Include/arm_nnfunctions.h:3836`
## arm_sub_scalar_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_sub_scalar_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise subtract of scalar and vector (scalar - vector)
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input scalar |
| input_2_vect | const int16_t * | in | pointer to input vector |
| input_1_offset | const int32_t | in | offset for input 1. Not used. |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Not used. |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int16_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Not used. |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3881`
## arm_elementwise_sub_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_sub_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise subtract of two vectors
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1. Not used. |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Not used. |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | input left shift. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int16_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Not used. |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:3923`
## arm_squared_difference_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_squared_difference_s8(
const int8_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int8_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)
```
s8 elementwise squared difference of two tensors with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input1_data | const int8_t * | in | pointer to input tensor 1 |
| input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
| input2_data | const int8_t * | in | pointer to input tensor 2 |
| input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
| input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input1_mult | const int32_t | in | multiplier for input 1 |
| input1_shift | const int32_t | in | shift for input 1 |
| input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input2_mult | const int32_t | in | multiplier for input 2 |
| input2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output_data | int8_t * | out | pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
Source: `Include/arm_nnfunctions.h:3969`
## arm_squared_difference_scalar_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_squared_difference_scalar_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s8 elementwise squared difference of scalar and vector.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int8_t * | in | pointer to input scalar |
| input_2_vect | const int8_t * | in | pointer to input vector |
| input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int8_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:4013`
## arm_elementwise_squared_difference_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_squared_difference_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s8 elementwise squared difference of two vectors.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int8_t * | in | pointer to input vector 1 |
| input_2_vect | const int8_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int8_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:4055`
## arm_squared_difference_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_squared_difference_s16(
const int16_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int16_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)
```
s16 elementwise squared difference of two tensors with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input1_data | const int16_t * | in | pointer to input tensor 1 |
| input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
| input2_data | const int16_t * | in | pointer to input tensor 2 |
| input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
| input1_offset | const int32_t | in | offset for input 1 |
| input1_mult | const int32_t | in | multiplier for input 1 |
| input1_shift | const int32_t | in | shift for input 1 |
| input2_offset | const int32_t | in | offset for input 2 |
| input2_mult | const int32_t | in | multiplier for input 2 |
| input2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output_data | int16_t * | out | pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
| out_offset | const int32_t | in | output offset |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
Source: `Include/arm_nnfunctions.h:4101`
## arm_squared_difference_scalar_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_squared_difference_scalar_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise squared difference of scalar and vector.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input scalar |
| input_2_vect | const int16_t * | in | pointer to input vector |
| input_1_offset | const int32_t | in | offset for input 1 |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2 |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int16_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:4145`
## arm_elementwise_squared_difference_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_squared_difference_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise squared difference of two vectors.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1 |
| input_1_mult | const int32_t | in | multiplier for input 1 |
| input_1_shift | const int32_t | in | shift for input 1 |
| input_2_offset | const int32_t | in | offset for input 2 |
| input_2_mult | const int32_t | in | multiplier for input 2 |
| input_2_shift | const int32_t | in | shift for input 2 |
| left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| output | int16_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:4187`
## arm_mul_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_mul_s8(
const int8_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int8_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input2_offset,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)
```
s8 elementwise multiplication of two tensors with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input1_data | const int8_t * | in | pointer to input tensor 1 |
| input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
| input2_data | const int8_t * | in | pointer to input tensor 2 |
| input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
| input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| output_data | int8_t * | out | pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
Source: `Include/arm_nnfunctions.h:4224`
## arm_mul_scalar_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_mul_scalar_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s8 elementwise multiplication of scalar and vector
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int8_t * | in | pointer to input scalar |
| input_2_vect | const int8_t * | in | pointer to input vector |
| input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| output | int8_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:4254`
## arm_elementwise_mul_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_mul_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s8 elementwise multiplication
Supported framework: TensorFlow Lite micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int8_t * | in | pointer to input vector 1 |
| input_2_vect | const int8_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
| input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
| output | int8_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Range: -128 to 127 |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:4283`
## arm_mul_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_mul_s16(
const int16_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int16_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input2_offset,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)
```
s16 elementwise multiplication of two tensors with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input1_data | const int16_t * | in | pointer to input tensor 1 |
| input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
| input2_data | const int16_t * | in | pointer to input tensor 2 |
| input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
| input1_offset | const int32_t | in | offset for input 1. Not used. |
| input2_offset | const int32_t | in | offset for input 2. Not used. |
| output_data | int16_t * | out | pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
| out_offset | const int32_t | in | output offset. Not used. |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
Source: `Include/arm_nnfunctions.h:4315`
## arm_mul_scalar_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_mul_scalar_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise multiplication of scalar and vector
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input scalar |
| input_2_vect | const int16_t * | in | pointer to input vector |
| input_1_offset | const int32_t | in | offset for input 1. Not used. |
| input_2_offset | const int32_t | in | offset for input 2. Not used. |
| output | int16_t * | out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Not used. |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:4345`
## arm_elementwise_mul_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_mul_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise multiplication
Supported framework: TensorFlow Lite micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1. Not used. |
| input_2_offset | const int32_t | in | offset for input 2. Not used. |
| output | int16_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Not used. |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:4374`
## arm_minimum_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_minimum_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)
```
s8 elementwise minimum w/ support for broadcasting and scalar inputs.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int8_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int8_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | int8_t * | out | Pointer to the output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions.h:4405`
## arm_maximum_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_maximum_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)
```
s8 elementwise maximum w/ support for broadcasting and scalar inputs.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int8_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int8_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | int8_t * | out | Pointer to the output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions.h:4432`
## arm_minimum_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_minimum_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
s16 elementwise minimum w/ support for broadcasting and scalar inputs.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int16_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int16_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | int16_t * | out | Pointer to the output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions.h:4459`
## arm_maximum_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_maximum_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
s16 elementwise maximum w/ support for broadcasting and scalar inputs.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int16_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int16_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | int16_t * | out | Pointer to the output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions.h:4486`
## arm_comparison_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_comparison_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
arm_nn_compare_operation operation
)
```
s8 elementwise comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int8_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int8_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| operation | arm_nn_compare_operation | in | Comparison operation to perform |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions.h:4527`
## arm_comparison_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_comparison_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
arm_nn_compare_operation operation
)
```
s16 elementwise comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int16_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int16_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
| operation | arm_nn_compare_operation | in | Comparison operation to perform |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
Source: `Include/arm_nnfunctions.h:4570`
## arm_equal_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_equal_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s8 elementwise equality comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int8_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int8_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4610`
## arm_not_equal_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_not_equal_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s8 elementwise inequality comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int8_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int8_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4649`
## arm_greater_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_greater_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s8 elementwise greater-than comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int8_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int8_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4688`
## arm_greater_equal_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_greater_equal_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s8 elementwise greater-or-equal comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int8_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int8_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4727`
## arm_less_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_less_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s8 elementwise less-than comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int8_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int8_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4766`
## arm_less_equal_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_less_equal_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s8 elementwise less-or-equal comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int8_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int8_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4805`
## arm_equal_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_equal_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s16 elementwise equality comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int16_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int16_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4844`
## arm_not_equal_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_not_equal_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s16 elementwise inequality comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int16_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int16_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4883`
## arm_greater_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_greater_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s16 elementwise greater-than comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int16_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int16_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4922`
## arm_greater_equal_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_greater_equal_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s16 elementwise greater-or-equal comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int16_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int16_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:4961`
## arm_less_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_less_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s16 elementwise less-than comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int16_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int16_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:5000`
## arm_less_equal_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_less_equal_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)
```
s16 elementwise less-or-equal comparison with support for broadcasting.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
| input_1_data | const int16_t * | in | Pointer to input1 tensor |
| input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
| input_2_data | const int16_t * | in | Pointer to input2 tensor |
| input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
| output_data | bool * | out | Pointer to the output tensor (bool values) |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| input_1_offset | const int32_t | in | Zero-point for input1 tensor |
| input_1_mult | const int32_t | in | Multiplier for input1 tensor |
| input_1_shift | const int32_t | in | Shift for input1 tensor |
| input_2_offset | const int32_t | in | Zero-point for input2 tensor |
| input_2_mult | const int32_t | in | Multiplier for input2 tensor |
| input_2_shift | const int32_t | in | Shift for input2 tensor |
| left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Source: `Include/arm_nnfunctions.h:5039`
## arm_relu_q7
`function` · `c`
```c
void arm_relu_q7(int8_t *data, uint16_t size)
```
Q7 RELU function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data | int8_t * | in, out | pointer to input |
| size | uint16_t | in | number of elements |
Source: `Include/arm_nnfunctions.h:5067`
## arm_relu6_q7
`function` · `c`
```c
void arm_relu6_q7(int8_t *data, uint16_t size)
```
Q7 RELU6 function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data | int8_t * | in, out | pointer to input |
| size | uint16_t | in | number of elements |
Source: `Include/arm_nnfunctions.h:5074`
## arm_relu_q15
`function` · `c`
```c
void arm_relu_q15(int16_t *data, uint16_t size)
```
Q15 RELU function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data | int16_t * | in, out | pointer to input |
| size | uint16_t | in | number of elements |
Source: `Include/arm_nnfunctions.h:5081`
## arm_clamp_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_clamp_s8(
const int8_t *input,
const int8_t act_min,
const int8_t act_max,
int8_t *output,
const int32_t output_size
)
```
S8 clamp function.
This function clamps each element in the input tensor to the range. This can be useful for activations such as relu(0, 6), relu(-1,1), etc.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input |
| act_min | const int8_t | in | Minimum value to clamp to |
| act_max | const int8_t | in | Maximum value to clamp to |
| output | int8_t * | out | Pointer to output |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:5095`
## arm_clamp_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_clamp_s16(
const int16_t *input,
const int16_t act_min,
const int16_t act_max,
int16_t *output,
const int32_t output_size
)
```
S16 clamp function.
This function clamps each element in the input tensor to the range. This can be useful for activations such as relu(0, 6), relu(-1,1), etc.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to input |
| act_min | const int16_t | in | Minimum value to clamp to |
| act_max | const int16_t | in | Maximum value to clamp to |
| output | int16_t * | out | Pointer to output |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:5113`
## arm_relu_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_relu_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
int8_t *output,
const int32_t output_size
)
```
S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier | const int32_t | in | Output multiplier |
| output_shift | const int32_t | in | Output shift |
| output | int8_t * | out | Pointer to the output buffer |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5132`
## arm_relu_generic_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_relu_generic_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
const int32_t act_min,
const int32_t act_max,
int8_t *output,
const int32_t output_size
)
```
S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier | const int32_t | in | Output multiplier |
| output_shift | const int32_t | in | Output shift |
| act_min | const int32_t | in | Minimum value to clamp the output to |
| act_max | const int32_t | in | Maximum value to clamp the output to |
| output | int8_t * | out | Pointer to the output buffer |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5155`
## arm_relu_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_relu_s16(
const int16_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
int16_t *output,
const int32_t output_size
)
```
S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier | const int32_t | in | Output multiplier |
| output_shift | const int32_t | in | Output shift |
| output | int16_t * | out | Pointer to the output buffer |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5178`
## arm_relu_generic_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_relu_generic_s16(
const int16_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
const int32_t act_min,
const int32_t act_max,
int16_t *output,
const int32_t output_size
)
```
S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier | const int32_t | in | Output multiplier |
| output_shift | const int32_t | in | Output shift |
| act_min | const int32_t | in | Minimum value to clamp the output to |
| act_max | const int32_t | in | Maximum value to clamp the output to |
| output | int16_t * | out | Pointer to the output buffer |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5201`
## arm_leaky_relu_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_leaky_relu_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
int8_t *output,
const int32_t output_size
)
```
S8 Leaky ReLU activation function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier_alpha | const int32_t | in | Output multiplier for the alpha parameter |
| output_shift_alpha | const int32_t | in | Output shift for the alpha parameter |
| output_multiplier_identity | const int32_t | in | Output multiplier for the identity parameter |
| output_shift_identity | const int32_t | in | Output shift for the identity parameter |
| output | int8_t * | out | Pointer to the output buffer |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5226`
## arm_leaky_relu_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_leaky_relu_s16(
const int16_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
int16_t *output,
const int32_t output_size
)
```
S16 Leaky ReLU activation function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier_alpha | const int32_t | in | Output multiplier for the alpha parameter |
| output_shift_alpha | const int32_t | in | Output shift for the alpha parameter |
| output_multiplier_identity | const int32_t | in | Output multiplier for the identity parameter |
| output_shift_identity | const int32_t | in | Output shift for the identity parameter |
| output | int16_t * | out | Pointer to the output buffer |
| output_size | const int32_t | in | Number of elements in the input tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5251`
## arm_logistic_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_logistic_s16(
const int16_t *input,
int16_t *output,
const int32_t input_size,
int32_t input_multiplier,
int32_t input_left_shift
)
```
Logistic activation function for s16.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input tensor |
| output | int16_t * | out | Pointer to the output tensor |
| input_size | const int32_t | in | Number of elements in the input tensor |
| input_multiplier | int32_t | in | Input quantization multiplier |
| input_left_shift | int32_t | in | Input quantization shift within the range [0, 31] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:5272`
## arm_tanh_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_tanh_s16(
const int16_t *input,
int16_t *output,
const int32_t input_size,
int32_t input_multiplier,
int32_t input_left_shift
)
```
Tanh activation function for s16.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input tensor |
| output | int16_t * | out | Pointer to the output tensor |
| input_size | const int32_t | in | Number of elements in the input tensor |
| input_multiplier | int32_t | in | Input quantization multiplier |
| input_left_shift | int32_t | in | Input quantization shift within the range [0, 31] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:5289`
## arm_nn_activation_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_activation_s16(
const int16_t *input,
int16_t *output,
const int32_t size,
const int32_t left_shift,
const arm_nn_activation_type type
)
```
s16 neural network activation function using direct table look-up
Supported framework: TensorFlow Lite for Microcontrollers. This activation function must be bit precise congruent with the corresponding TFLM tanh and sigmoid activation functions
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | pointer to input data |
| output | int16_t * | out | pointer to output |
| size | const int32_t | in | number of elements |
| left_shift | const int32_t | in | bit-width of the integer part, assumed to be smaller than 3. |
| type | const arm_nn_activation_type | in | type of activation functions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:5308`
## arm_hard_swish_compat_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_hard_swish_compat_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier_fp,
const int32_t output_multiplier_exp,
const int32_t relu_multiplier_fp,
const int32_t relu_multiplier_exp,
int8_t *output,
const int32_t output_size
)
```
S8 Hard-Swish activation function (compatibility version).
This version is compatible with TFLite implementation of Hard-Swish. hires_input_scale = (1.0 / 128.0) * float(input_scale) relu_scale = 3.0 / 32768.0 out_mul_real = hires_input_scale / float(output_scale) relu_mul_real = hires_input_scale / relu_scale output_multiplier_fp, output_multiplier_exp = to_q15_exp(out_mul_real) relu_multiplier_fp, relu_multiplier_exp = to_q15_exp(relu_mul_real) Here to_q15_exp quantizes to Q31 with a frexp exponent, then rounds and saturates the Q31 multiplier to Q15. For input_scale = output_scale = 0.125, the output pair is (16384, -6) and the ReLU pair is (21845, 4).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier_fp | const int32_t | in | Output multiplier in fixed point format |
| output_multiplier_exp | const int32_t | in | Exponent for output multiplier |
| relu_multiplier_fp | const int32_t | in | ReLU6 multiplier in fixed point format |
| relu_multiplier_exp | const int32_t | in | Exponent for ReLU6 multiplier |
| output | int8_t * | out | Pointer to the output buffer |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR if output_multiplier_exp is positive. |
Source: `Include/arm_nnfunctions.h:5338`
## arm_hard_swish_precise_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_hard_swish_precise_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
const int32_t relu_q3,
const int32_t relu_q6,
const int32_t prescale,
int8_t *output,
const int32_t output_size
)
```
S8 Hard-Swish activation function (precise version).
This version uses int32_t for intermediate computations to provide better accuracy. relu_q3, relu_q6 = round(3 / input_scale), round(6 / input_scale) prescale = min value to multiply input by to avoid overflow in intermediate computations M = ((input_scale**2) / (6.0 * output_scale)) * (1 << prescale) output_multiplier, output_shift = quantize_multiplier(M)
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier | const int32_t | in | Output multiplier |
| output_shift | const int32_t | in | Output shift |
| relu_q3 | const int32_t | in | ReLU6 Q3 value |
| relu_q6 | const int32_t | in | ReLU6 Q6 value |
| prescale | const int32_t | in | Prescale to apply to input |
| output | int8_t * | out | Pointer to the output buffer |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5369`
## arm_hard_swish_precise_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_hard_swish_precise_s16(
const int16_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
const int32_t relu_q3,
const int32_t relu_q6,
const int32_t prescale,
int16_t *output,
const int32_t output_size
)
```
S16 Hard-Swish activation function (precise version).
This version uses int32_t for intermediate computations to provide better accuracy. relu_q3, relu_q6 = round(3 / input_scale), round(6 / input_scale) prescale = min value to multiply input by to avoid overflow in intermediate computations M = ((input_scale**2) / (6.0 * output_scale)) * (1 << prescale) output_multiplier, output_shift = quantize_multiplier(M)
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier | const int32_t | in | Output multiplier |
| output_shift | const int32_t | in | Output shift |
| relu_q3 | const int32_t | in | ReLU6 Q3 value |
| relu_q6 | const int32_t | in | ReLU6 Q6 value |
| prescale | const int32_t | in | Prescale to apply to input |
| output | int16_t * | out | Pointer to the output buffer |
| output_size | const int32_t | in | Number of elements in the tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5401`
## arm_prelu_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_prelu_s8(
const cmsis_nn_dims *input_dims,
const int8_t *input,
const cmsis_nn_dims *alpha_dims,
const int8_t *alpha,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t output_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
const cmsis_nn_dims *output_dims,
int8_t *output
)
```
S8 PReLU activation function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input | const int8_t * | in | Pointer to the input buffer |
| alpha_dims | const cmsis_nn_dims * | in | Alpha tensor dimensions. Format: [N, H, W, C] |
| alpha | const int8_t * | in | Pointer to the alpha buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| alpha_offset | const int32_t | in | Alpha tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier_identity | const int32_t | in | Output multiplier 1 |
| output_shift_identity | const int32_t | in | Output shift 1 |
| output_multiplier_alpha | const int32_t | in | Output multiplier 2 |
| output_shift_alpha | const int32_t | in | Output shift 2 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output | int8_t * | out | Pointer to the output buffer |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, alpha does not broadcast into the input, or the output dimensions do not match the input dimensions. |
Source: `Include/arm_nnfunctions.h:5432`
## arm_elementwise_prelu_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_prelu_s8(
const int8_t *input,
const int8_t *alpha,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t out_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
int8_t *output,
const int32_t block_size
)
```
Elementwise S8 PReLU activation function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input buffer |
| alpha | const int8_t * | in | Pointer to the alpha buffer (same shape as input) |
| input_offset | const int32_t | in | Input tensor zero offset |
| alpha_offset | const int32_t | in | Alpha tensor zero offset |
| out_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
| output_shift_identity | const int32_t | in | Output shift when input >= 0 |
| output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
| output_shift_alpha | const int32_t | in | Output shift when input < 0 |
| output | int8_t * | out | Pointer to the output buffer |
| block_size | const int32_t | in | Number of elements to process |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5462`
## arm_prelu_scalar_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_prelu_scalar_s8(
const int8_t *scalar_vect,
const int8_t *non_scalar_vect,
const bool scalar_is_input,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t output_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
int8_t *output,
const int32_t block_size
)
```
Scalar S8 PReLU activation function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| scalar_vect | const int8_t * | in | Pointer to the scalar buffer (single value) |
| non_scalar_vect | const int8_t * | in | Pointer to the non-scalar buffer |
| scalar_is_input | const bool | in | True if the scalar buffer holds the input value, false if it holds alpha |
| input_offset | const int32_t | in | Input tensor zero offset |
| alpha_offset | const int32_t | in | Alpha tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
| output_shift_identity | const int32_t | in | Output shift when input >= 0 |
| output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
| output_shift_alpha | const int32_t | in | Output shift when input < 0 |
| output | int8_t * | out | Pointer to the output buffer |
| block_size | const int32_t | in | Number of elements to process when the non-scalar vector is used |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_MATH_SUCCESS |
Source: `Include/arm_nnfunctions.h:5491`
## arm_prelu_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_prelu_s16(
const cmsis_nn_dims *input_dims,
const int16_t *input,
const cmsis_nn_dims *alpha_dims,
const int16_t *alpha,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t output_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
const cmsis_nn_dims *output_dims,
int16_t *output
)
```
S16 PReLU activation function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| input | const int16_t * | in | Pointer to the input buffer |
| alpha_dims | const cmsis_nn_dims * | in | Alpha tensor dimensions. Format: [N, H, W, C] |
| alpha | const int16_t * | in | Pointer to the alpha buffer |
| input_offset | const int32_t | in | Input tensor zero offset |
| alpha_offset | const int32_t | in | Alpha tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
| output_shift_identity | const int32_t | in | Output shift when input >= 0 |
| output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
| output_shift_alpha | const int32_t | in | Output shift when input < 0 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| output | int16_t * | out | Pointer to the output buffer |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, alpha does not broadcast into the input, or the output dimensions do not match the input dimensions. |
Source: `Include/arm_nnfunctions.h:5524`
## arm_elementwise_prelu_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_prelu_s16(
const int16_t *input,
const int16_t *alpha,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t out_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
int16_t *output,
const int32_t block_size
)
```
Elementwise S16 PReLU activation function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input buffer |
| alpha | const int16_t * | in | Pointer to the alpha buffer (same shape as input) |
| input_offset | const int32_t | in | Input tensor zero offset |
| alpha_offset | const int32_t | in | Alpha tensor zero offset |
| out_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
| output_shift_identity | const int32_t | in | Output shift when input >= 0 |
| output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
| output_shift_alpha | const int32_t | in | Output shift when input < 0 |
| output | int16_t * | out | Pointer to the output buffer |
| block_size | const int32_t | in | Number of elements to process |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:5554`
## arm_prelu_scalar_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_prelu_scalar_s16(
const int16_t *scalar_vect,
const int16_t *non_scalar_vect,
const bool scalar_is_input,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t output_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
int16_t *output,
const int32_t block_size
)
```
Scalar S16 PReLU activation function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| scalar_vect | const int16_t * | in | Pointer to the scalar buffer (single value) |
| non_scalar_vect | const int16_t * | in | Pointer to the non-scalar buffer |
| scalar_is_input | const bool | in | True if the scalar buffer holds the input value, false if it holds alpha |
| input_offset | const int32_t | in | Input tensor zero offset |
| alpha_offset | const int32_t | in | Alpha tensor zero offset |
| output_offset | const int32_t | in | Output tensor zero offset |
| output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
| output_shift_identity | const int32_t | in | Output shift when input >= 0 |
| output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
| output_shift_alpha | const int32_t | in | Output shift when input < 0 |
| output | int16_t * | out | Pointer to the output buffer |
| block_size | const int32_t | in | Number of elements to process when the non-scalar vector is used |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnfunctions.h:5583`
## arm_avgpool_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_avgpool_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params *pool_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 average pooling function.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context. Size ctx->buf with arm_avgpool_s8_get_buffer_size(output_dims->w, input_dims->c). Note that it takes the output width and the input channel count rather than a `cmsis_nn_dims`. `arm_avgpool_s8_get_buffer_size_dsp()` and `arm_avgpool_s8_get_buffer_size_mve()` size the same buffer for a specific target. It returns 0 where no buffer is needed, in which case ctx->buf may be NULL. The caller is expected to clear the buffer, if applicable, for security reasons. |
| pool_params | const cmsis_nn_pool_params * | in | Pooling parameters |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [H, W] Argument N and C are not used. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN. |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns (an extent of 0 or less), which writes nothing and does not use ctx `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments, including a negative channel count, a pooling window that does not overlap the input, window positions (output index times stride minus padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t, or, on builds without MVE, a NULL ctx, or a NULL ctx->buf where the sizer asks for a buffer. Nothing is written to output_data then. |
Source: `Include/arm_nnfunctions.h:5638`
## arm_avgpool_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_avgpool_s8_get_buffer_size(const int dim_dst_width, const int ch_src)
```
Get the required buffer size for S8 average pooling function.
Unlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. `arm_avgpool_s8()` depends on that sentinel being non-zero, since it reads a non-zero size as "ctx->buf is required" before touching the accumulator buffer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dim_dst_width | const int | in | output tensor dimension |
| ch_src | const int | in | number of input tensor channels |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:5659`
## arm_avgpool_s8_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_avgpool_s8_get_buffer_size_dsp(const int dim_dst_width, const int ch_src)
```
Get the required buffer size for S8 average pooling function for processors with DSP extension.
Unlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. `arm_avgpool_s8()` depends on that sentinel being non-zero, since it reads a non-zero size as "ctx->buf is required" before touching the accumulator buffer.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_avgpool_s8_get_buffer_size()`.
:::
:::note
This is the leg that computes a byte count, so it also validates ch_src like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dim_dst_width | const int | in | output tensor dimension |
| ch_src | const int | in | number of input tensor channels |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:5671`
## arm_avgpool_s8_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_avgpool_s8_get_buffer_size_mve(const int dim_dst_width, const int ch_src)
```
Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case.
Unlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. `arm_avgpool_s8()` depends on that sentinel being non-zero, since it reads a non-zero size as "ctx->buf is required" before touching the accumulator buffer.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_avgpool_s8_get_buffer_size()`.
:::
:::note
This variant needs no buffer, so it returns 0 for every in-range shape. It still validates ch_src like the top-level dispatcher and the DSP leg, returning -1 for a negative ch_src or one whose byte count would not fit in an int32_t, so all three entry points answer an out-of-range shape alike.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dim_dst_width | const int | in | output tensor dimension |
| ch_src | const int | in | number of input tensor channels |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:5684`
## arm_avgpool_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_avgpool_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params *pool_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)
```
s16 average pooling function.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context. Size ctx->buf with arm_avgpool_s16_get_buffer_size(output_dims->w, input_dims->c). Note that it takes the output width and the input channel count rather than a `cmsis_nn_dims`. `arm_avgpool_s16_get_buffer_size_dsp()` and `arm_avgpool_s16_get_buffer_size_mve()` size the same buffer for a specific target. It returns 0 where no buffer is needed, in which case ctx->buf may be NULL. The caller is expected to clear the buffer, if applicable, for security reasons. |
| pool_params | const cmsis_nn_pool_params * | in | Pooling parameters |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [H, W] Argument N and C are not used. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN. |
| output_data | int16_t * | out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns (an extent of 0 or less), which writes nothing and does not use ctx `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments, including a negative channel count, a pooling window that does not overlap the input, window positions (output index times stride minus padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t, or, on builds that use the buffer, a NULL ctx, or a NULL ctx->buf where the sizer asks for a buffer. Nothing is written to output_data then. |
Source: `Include/arm_nnfunctions.h:5721`
## arm_avgpool_s16_get_buffer_size
`function` · `c`
```c
int32_t arm_avgpool_s16_get_buffer_size(const int dim_dst_width, const int ch_src)
```
Get the required buffer size for S16 average pooling function.
As in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dim_dst_width | const int | in | output tensor dimension |
| ch_src | const int | in | number of input tensor channels |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:5741`
## arm_avgpool_s16_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_avgpool_s16_get_buffer_size_dsp(const int dim_dst_width, const int ch_src)
```
Get the required buffer size for S16 average pooling function for processors with DSP extension.
As in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_avgpool_s16_get_buffer_size()`.
:::
:::note
This is the leg that computes a byte count, so it also validates ch_src like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dim_dst_width | const int | in | output tensor dimension |
| ch_src | const int | in | number of input tensor channels |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:5753`
## arm_avgpool_s16_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_avgpool_s16_get_buffer_size_mve(const int dim_dst_width, const int ch_src)
```
Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case.
As in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_avgpool_s16_get_buffer_size()`.
:::
:::note
This variant needs no buffer, so it returns 0 for every in-range shape. It still validates ch_src like the top-level dispatcher and the DSP leg, returning -1 for a negative ch_src or one whose byte count would not fit in an int32_t, so all three entry points answer an out-of-range shape alike.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dim_dst_width | const int | in | output tensor dimension |
| ch_src | const int | in | number of input tensor channels |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:5766`
## arm_max_pool_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_max_pool_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params *pool_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 max pooling function.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_max_pool_s8_get_buffer_size(). Max pooling needs no accumulator scratch, unlike `arm_avgpool_s8()`, whose sizer does not describe this argument. |
| pool_params | const cmsis_nn_pool_params * | in | Pooling parameters |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] |
| input_data | const int8_t * | in | Input (activation) data pointer. The input tensor must not overlap with the output tensor. Data type: int8 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [H, W] Argument N and C are not used. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN. |
| output_data | int8_t * | out | Output data pointer. Data type: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns, which writes nothing `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments: a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to output_data then. |
Source: `Include/arm_nnfunctions.h:5798`
## arm_max_pool_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_max_pool_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params *pool_params,
const cmsis_nn_dims *input_dims,
const int16_t *src,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int16_t *dst
)
```
s16 max pooling function.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_max_pool_s16_get_buffer_size(). Max pooling needs no accumulator scratch, unlike `arm_avgpool_s16()`, whose sizer does not describe this argument. |
| pool_params | const cmsis_nn_pool_params * | in | Pooling parameters |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] |
| src | const int16_t * | in | Input (activation) data pointer. The input tensor must not overlap with the output tensor. Data type: int16 |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [H, W] Argument N and C are not used. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN. |
| dst | int16_t * | in, out | Output data pointer. Data type: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns, which writes nothing `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments: a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then. |
Source: `Include/arm_nnfunctions.h:5836`
## arm_softmax_s8
`function` · `c`
```c
void arm_softmax_s8(
const int8_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const int32_t diff_min,
int8_t *output
)
```
S8 softmax function.
:::note
Supported framework: TensorFlow Lite micro (bit-accurate)
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input tensor |
| num_rows | const int32_t | in | Number of rows in the input tensor |
| row_size | const int32_t | in | Number of elements in each input row |
| mult | const int32_t | in | Input quantization multiplier |
| shift | const int32_t | in | Input quantization shift within the range [0, 31] |
| diff_min | const int32_t | in | Minimum difference with max in row. Used to check if the quantized exponential operation can be performed |
| output | int8_t * | out | Pointer to the output tensor |
Source: `Include/arm_nnfunctions.h:5864`
## arm_softmax_s8_s16
`function` · `c`
```c
void arm_softmax_s8_s16(
const int8_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const int32_t diff_min,
int16_t *output
)
```
S8 to s16 softmax function.
:::note
Supported framework: TensorFlow Lite micro (bit-accurate)
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input tensor |
| num_rows | const int32_t | in | Number of rows in the input tensor |
| row_size | const int32_t | in | Number of elements in each input row |
| mult | const int32_t | in | Input quantization multiplier |
| shift | const int32_t | in | Input quantization shift within the range [0, 31] |
| diff_min | const int32_t | in | Minimum difference with max in row. Used to check if the quantized exponential operation can be performed |
| output | int16_t * | out | Pointer to the output tensor |
Source: `Include/arm_nnfunctions.h:5886`
## arm_softmax_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_softmax_s16(
const int16_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const cmsis_nn_softmax_lut_s16 *softmax_params,
int16_t *output
)
```
S16 softmax function.
:::note
Supported framework: TensorFlow Lite micro (bit-accurate)
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input tensor |
| num_rows | const int32_t | in | Number of rows in the input tensor |
| row_size | const int32_t | in | Number of elements in each input row |
| mult | const int32_t | in | Input quantization multiplier |
| shift | const int32_t | in | Input quantization shift within the range [0, 31] |
| softmax_params | const cmsis_nn_softmax_lut_s16 * | in | Softmax s16 layer parameters with two pointers to LUTs speficied below. For indexing the high 9 bits are used and 7 remaining for interpolation. That means 512 entries for the 9-bit indexing and 1 extra for interpolation, i.e. 513 values for each LUT.
- Lookup table for exp(x), where x uniform distributed between [-10.0 , 0.0] - Lookup table for 1 / (1 + x), where x uniform distributed between [0.0 , 1.0] |
| output | int16_t * | out | Pointer to the output tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Source: `Include/arm_nnfunctions.h:5915`
## arm_softmax_u8
`function` · `c`
```c
void arm_softmax_u8(
const uint8_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const int32_t diff_min,
uint8_t *output
)
```
U8 softmax function.
:::note
Supported framework: TensorFlow Lite micro (bit-accurate)
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const uint8_t * | in | Pointer to the input tensor |
| num_rows | const int32_t | in | Number of rows in the input tensor |
| row_size | const int32_t | in | Number of elements in each input row |
| mult | const int32_t | in | Input quantization multiplier |
| shift | const int32_t | in | Input quantization shift within the range [0, 31] |
| diff_min | const int32_t | in | Minimum difference with max in row. Used to check if the quantized exponential operation can be performed |
| output | uint8_t * | out | Pointer to the output tensor |
Source: `Include/arm_nnfunctions.h:5938`
## arm_reshape_s8
`function` · `c`
```c
void arm_reshape_s8(const int8_t *input, int8_t *output, const uint32_t total_size)
```
Reshape a s8 vector into another with different shape.
:::note
The output is expected to be in a memory area that does not overlap with the input's
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | points to the s8 input vector |
| output | int8_t * | out | points to the s8 output vector |
| total_size | const uint32_t | in | total size of the input and output vectors in bytes |
Source: `Include/arm_nnfunctions.h:5960`
## arm_resize_nearest_neighbor_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_resize_nearest_neighbor_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_resize_params *resize_params,
const cmsis_nn_dims *input_shape,
const int8_t *input_data,
const cmsis_nn_dims *output_size_shape,
const int32_t *output_size_data,
const cmsis_nn_dims *output_shape,
int8_t *output_data
)
```
Nearest neighbor resize function for s8 data.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Pointer to the context buffer. The buffer must hold at least (output_height + output_width) int32_t elements. |
| resize_params | const cmsis_nn_resize_params * | in | Resize parameters |
| input_shape | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C] |
| input_data | const int8_t * | in | Pointer to input tensor data |
| output_size_shape | const cmsis_nn_dims * | in | Output size tensor dimensions |
| output_size_data | const int32_t * | in | Output size tensor data |
| output_shape | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C] |
| output_data | int8_t * | out | Pointer to output tensor data |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:5982`
## arm_resize_nearest_neighbor_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_resize_nearest_neighbor_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_resize_params *resize_params,
const cmsis_nn_dims *input_shape,
const int16_t *input_data,
const cmsis_nn_dims *output_size_shape,
const int32_t *output_size_data,
const cmsis_nn_dims *output_shape,
int16_t *output_data
)
```
Nearest neighbor resize function for s16 data.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Pointer to the context buffer. The buffer must hold at least (output_height + output_width) int32_t elements. |
| resize_params | const cmsis_nn_resize_params * | in | Resize parameters |
| input_shape | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C] |
| input_data | const int16_t * | in | Pointer to input tensor data |
| output_size_shape | const cmsis_nn_dims * | in | Output size tensor dimensions |
| output_size_data | const int32_t * | in | Output size tensor data |
| output_shape | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C] |
| output_data | int16_t * | out | Pointer to output tensor data |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6011`
## arm_space_to_depth_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_space_to_depth_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t block_size,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Space to Depth function for s8 data type.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to the input tensor. Data type: int8 |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| block_size | const int32_t | in | Block size for space to depth transformation |
| output_data | int8_t * | out | Pointer to the output tensor. Data type: int8 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H/block_size, W/block_size, C_IN*block_size*block_size] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6036`
## arm_space_to_depth_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_space_to_depth_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t block_size,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Space to Depth function for s16 data type.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to the input tensor. Data type: int16 |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| block_size | const int32_t | in | Block size for space to depth transformation |
| output_data | int16_t * | out | Pointer to the output tensor. Data type: int16 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H/block_size, W/block_size, C_IN*block_size*block_size] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6058`
## arm_depth_to_space_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_depth_to_space_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t block_size,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Depth to Space function for s8 data type.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to the input tensor. Data type: int8 |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| block_size | const int32_t | in | Block size for depth to space transformation |
| output_data | int8_t * | out | Pointer to the output tensor. Data type: int8 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H*block_size, W*block_size, C_IN/(block_size*block_size)] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6080`
## arm_depth_to_space_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_depth_to_space_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t block_size,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Depth to Space function for s16 data type.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to the input tensor. Data type: int16 |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| block_size | const int32_t | in | Block size for depth to space transformation |
| output_data | int16_t * | out | Pointer to the output tensor. Data type: int16 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H*block_size, W*block_size, C_IN/(block_size*block_size)] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6102`
## arm_space_to_batch_nd_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_space_to_batch_nd_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_tile *block_shape,
const cmsis_nn_dims *pad,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t output_offset
)
```
Space to Batch ND function for s8 data type.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to the input tensor. Data type: int8 |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| block_shape | const cmsis_nn_tile * | in | Block shape for space to batch transformation |
| pad | const cmsis_nn_dims * | in | Padding for height and width. Format: [n->top, h->left, w->bottom, c->right] |
| output_data | int8_t * | out | Pointer to the output tensor. Data type: int8 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N*block_shape[0]*block_shape[1], (H + pad_top + pad_bottom)/block_shape[0], (W + pad_left + pad_right)/block_shape[1], C_IN] |
| output_offset | const int32_t | in | Zero offset for the output tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6126`
## arm_space_to_batch_nd_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_space_to_batch_nd_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_tile *block_shape,
const cmsis_nn_dims *pad,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t output_offset
)
```
Space to Batch ND function for s16 data type.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to the input tensor. Data type: int16 |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| block_shape | const cmsis_nn_tile * | in | Block shape for space to batch transformation |
| pad | const cmsis_nn_dims * | in | Padding for height and width. Format: [n->top, h->left, w->bottom, c->right] |
| output_data | int16_t * | out | Pointer to the output tensor. Data type: int16 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N*block_shape[0]*block_shape[1], (H + pad_top + pad_bottom)/block_shape[0], (W + pad_left + pad_right)/block_shape[1], C_IN] |
| output_offset | const int32_t | in | Zero offset for the output tensor. NOT USED. Assume symmetric quantization for s16. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6152`
## arm_batch_to_space_nd_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_batch_to_space_nd_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_tile *block_shape,
const cmsis_nn_dims *crop,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Batch to Space ND function for s8 data type.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to the input tensor. Data type: int8 |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| block_shape | const cmsis_nn_tile * | in | Block shape for batch to space transformation |
| crop | const cmsis_nn_dims * | in | Cropping for height and width. Format: [n->top, h->left, w->bottom, c->right] |
| output_data | int8_t * | out | Pointer to the output tensor. Data type: int8 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N/(block_shape[0]*block_shape[1]), H*block_shape[0] - crop_top - crop_bottom, W*block_shape[1] - crop_left - crop_right, C_IN] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6177`
## arm_batch_to_space_nd_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_batch_to_space_nd_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_tile *block_shape,
const cmsis_nn_dims *crop,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Batch to Space ND function for s16 data type.
- Supported Framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to the input tensor. Data type: int16 |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| block_shape | const cmsis_nn_tile * | in | Block shape for batch to space transformation |
| crop | const cmsis_nn_dims * | in | Cropping for height and width. Format: [n->top, h->left, w->bottom, c->right] |
| output_data | int16_t * | out | Pointer to the output tensor. Data type: int16 |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N/(block_shape[0]*block_shape[1]), H*block_shape[0] - crop_top - crop_bottom, W*block_shape[1] - crop_left - crop_right, C_IN] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6201`
## arm_transpose_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_s8(
const int8_t *input_data,
int8_t *const output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const output_dims,
const cmsis_nn_transpose_params *const transpose_params
)
```
Basic transpose function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
| output_data | int8_t *const | out | Output data pointer. Data type: int8 |
| input_dims | const cmsis_nn_dims *const | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions. Format may be arbitrary relative to input format. The output dimension will depend on the permutation dimensions. In other words the out dimensions are the result of applying the permutation to the input dimensions. The first transpose_params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[permutations[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not. |
| transpose_params | const cmsis_nn_transpose_params *const | in | Transpose parameters. Contains permutation dimensions. num_dims must be in [1, 4] and permutations must be a bijection over [0, num_dims - 1]. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6235`
## arm_transpose_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_transpose_s16(
const int16_t *input_data,
int16_t *const output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const output_dims,
const cmsis_nn_transpose_params *const transpose_params
)
```
Basic s16 transpose function.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
| output_data | int16_t *const | out | Output data pointer. Data type: int16 |
| input_dims | const cmsis_nn_dims *const | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
| output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions. Format may be arbitrary relative to input format. The output dimension will depend on the permutation dimensions. In other words the out dimensions are the result of applying the permutation to the input dimensions. The first transpose_params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[permutations[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not. |
| transpose_params | const cmsis_nn_transpose_params *const | in | Transpose parameters. Contains permutation dimensions. num_dims must be in [1, 4] and permutations must be a bijection over [0, num_dims - 1]. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6263`
## arm_concatenation_s8_x
`function` · `c`
```c
void arm_concatenation_s8_x(
const int8_t *input,
const uint16_t input_x,
const uint16_t input_y,
const uint16_t input_z,
const uint16_t input_w,
int8_t *output,
const uint16_t output_x,
const uint32_t offset_x
)
```
int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate. The argument offset_x will be used to store the input tensor in the correct position in the output tensor
i.e. offset_x = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_x(&input[i], ..., &output, ..., ..., offset_x) offset_x += input_x[i] }
This function assumes that the output tensor has:
1. The same height of the input tensor
2. The same number of channels of the input tensor
3. The same batch size of the input tensor
Unless specified otherwise, arguments are mandatory.
:::note
This function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation
:::
**Input constraints** offset_x is less than output_x
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input tensor. Input tensor must not overlap with the output tensor. |
| input_x | const uint16_t | in | Width of input tensor |
| input_y | const uint16_t | in | Height of input tensor |
| input_z | const uint16_t | in | Channels in input tensor |
| input_w | const uint16_t | in | Batch size in input tensor |
| output | int8_t * | out | Pointer to output tensor. Expected to be at least (input_x * input_y * input_z * input_w) + offset_x bytes. |
| output_x | const uint16_t | in | Width of output tensor |
| offset_x | const uint32_t | in | The offset (in number of elements) on the X axis to start concatenating the input tensor It is user responsibility to provide the correct value |
Source: `Include/arm_nnfunctions.h:6312`
## arm_concatenation_s8_y
`function` · `c`
```c
void arm_concatenation_s8_y(
const int8_t *input,
const uint16_t input_x,
const uint16_t input_y,
const uint16_t input_z,
const uint16_t input_w,
int8_t *output,
const uint16_t output_y,
const uint32_t offset_y
)
```
int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate. The argument offset_y will be used to store the input tensor in the correct position in the output tensor
i.e. offset_y = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_y(&input[i], ..., &output, ..., ..., offset_y) offset_y += input_y[i] }
This function assumes that the output tensor has:
1. The same width of the input tensor
2. The same number of channels of the input tensor
3. The same batch size of the input tensor
Unless specified otherwise, arguments are mandatory.
:::note
This function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation
:::
**Input constraints** offset_y is less than output_y
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input tensor. Input tensor must not overlap with the output tensor. |
| input_x | const uint16_t | in | Width of input tensor |
| input_y | const uint16_t | in | Height of input tensor |
| input_z | const uint16_t | in | Channels in input tensor |
| input_w | const uint16_t | in | Batch size in input tensor |
| output | int8_t * | out | Pointer to output tensor. Expected to be at least (input_z * input_w * input_x * input_y) + offset_y bytes. |
| output_y | const uint16_t | in | Height of output tensor |
| offset_y | const uint32_t | in | The offset on the Y axis to start concatenating the input tensor It is user responsibility to provide the correct value |
Source: `Include/arm_nnfunctions.h:6359`
## arm_concatenation_s8_z
`function` · `c`
```c
void arm_concatenation_s8_z(
const int8_t *input,
const uint16_t input_x,
const uint16_t input_y,
const uint16_t input_z,
const uint16_t input_w,
int8_t *output,
const uint16_t output_z,
const uint32_t offset_z
)
```
int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate. The argument offset_z will be used to store the input tensor in the correct position in the output tensor
i.e. offset_z = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_z(&input[i], ..., &output, ..., ..., offset_z) offset_z += input_z[i] }
This function assumes that the output tensor has:
1. The same width of the input tensor
2. The same height of the input tensor
3. The same batch size of the input tensor
Unless specified otherwise, arguments are mandatory.
:::note
This function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation
:::
**Input constraints** offset_z is less than output_z
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input tensor. Input tensor must not overlap with output tensor. |
| input_x | const uint16_t | in | Width of input tensor |
| input_y | const uint16_t | in | Height of input tensor |
| input_z | const uint16_t | in | Channels in input tensor |
| input_w | const uint16_t | in | Batch size in input tensor |
| output | int8_t * | out | Pointer to output tensor. Expected to be at least (input_x * input_y * input_z * input_w) + offset_z bytes. |
| output_z | const uint16_t | in | Channels in output tensor |
| offset_z | const uint32_t | in | The offset on the Z axis to start concatenating the input tensor It is user responsibility to provide the correct value |
Source: `Include/arm_nnfunctions.h:6406`
## arm_concatenation_s8_w
`function` · `c`
```c
void arm_concatenation_s8_w(
const int8_t *input,
const uint16_t input_x,
const uint16_t input_y,
const uint16_t input_z,
const uint16_t input_w,
int8_t *output,
const uint32_t offset_w
)
```
int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to concatenate. The argument offset_w will be used to store the input tensor in the correct position in the output tensor
i.e. offset_w = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_w(&input[i], ..., &output, ..., ..., offset_w) offset_w += input_w[i] }
This function assumes that the output tensor has:
1. The same width of the input tensor
2. The same height of the input tensor
3. The same number o channels of the input tensor
Unless specified otherwise, arguments are mandatory.
:::note
This function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input tensor |
| input_x | const uint16_t | in | Width of input tensor |
| input_y | const uint16_t | in | Height of input tensor |
| input_z | const uint16_t | in | Channels in input tensor |
| input_w | const uint16_t | in | Batch size in input tensor |
| output | int8_t * | out | Pointer to output tensor. Expected to be at least input_x * input_y * input_z * input_w bytes. |
| offset_w | const uint32_t | in | The offset on the W axis to start concatenating the input tensor It is user responsibility to provide the correct value |
Source: `Include/arm_nnfunctions.h:6449`
## arm_concatenation_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_concatenation_s8(
const int8_t *const *input_data,
const int32_t inputs_count,
const int32_t *input_concat_dims,
const int32_t axis,
int8_t *output_data,
const int32_t output_dims,
const int32_t *output_shape
)
```
int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis
:::note
This function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t *const * | in | Pointer to input tensors |
| inputs_count | const int32_t | in | Number of input tensors |
| input_concat_dims | const int32_t * | in | Dimensions of the input tensors along the target axis |
| axis | const int32_t | in | Target axis to concatenate the input tensors |
| output_data | int8_t * | out | Pointer to output tensor |
| output_dims | const int32_t | in | Output tensor dimensions |
| output_shape | const int32_t * | in | Output tensor shape |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:6474`
## arm_concatenation_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_concatenation_s16(
const int16_t *const *input_data,
const int32_t inputs_count,
const int32_t *input_concat_dims,
const int32_t axis,
int16_t *output_data,
const int32_t output_dims,
const int32_t *output_shape
)
```
int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis
:::note
This function, data layout independent, can be used to concatenate either int16 or uint16 tensors because it does not involve any arithmetic operation
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t *const * | in | Pointer to input tensors |
| inputs_count | const int32_t | in | Number of input tensors |
| input_concat_dims | const int32_t * | in | Dimensions of the input tensors along the target axis |
| axis | const int32_t | in | Target axis to concatenate the input tensors |
| output_data | int16_t * | out | Pointer to output tensor |
| output_dims | const int32_t | in | Output tensor dimensions |
| output_shape | const int32_t * | in | Output tensor shape |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:6499`
## arm_concatenation_s32
`function` · `c`
```c
arm_cmsis_nn_status arm_concatenation_s32(
const int32_t *const *input_data,
const int32_t inputs_count,
const int32_t *input_concat_dims,
const int32_t axis,
int32_t *output_data,
const int32_t output_dims,
const int32_t *output_shape
)
```
int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis
:::note
This function, data layout independent, can be used to concatenate either int32 or uint32 tensors because it does not involve any arithmetic operation
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int32_t *const * | in | Pointer to input tensors |
| inputs_count | const int32_t | in | Number of input tensors |
| input_concat_dims | const int32_t * | in | Dimensions of the input tensors along the target axis |
| axis | const int32_t | in | Target axis to concatenate the input tensors |
| output_data | int32_t * | out | Pointer to output tensor |
| output_dims | const int32_t | in | Output tensor dimensions |
| output_shape | const int32_t * | in | Output tensor shape |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:6524`
## arm_split_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_split_s8(
const int8_t *input_data,
const int32_t input_dims,
const int32_t *input_shape,
const int32_t axis,
const int32_t num_splits,
const int32_t *split_dims,
int8_t *const *output_data
)
```
int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis
:::note
This function, data layout independent, can be used to split either int8 or uint8 tensors because it does not involve any arithmetic operation.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to the flattened input tensor data. |
| input_dims | const int32_t | in | Number of dimensions in input_shape. |
| input_shape | const int32_t * | in | Array of length input_dims describing the shape of input_data. |
| axis | const int32_t | in | Axis along which to split (0 <= axis < input_dims). |
| num_splits | const int32_t | in | Number of output tensors to produce. |
| split_dims | const int32_t * | in | Array of length num_splits giving size of each slice along axis. |
| output_data | int8_t *const * | out | Array of pointers; output_data[i] points to storage for the i-th output tensor. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR if split_dims sum mismatch. |
Source: `Include/arm_nnfunctions.h:6547`
## arm_split_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_split_s16(
const int16_t *input_data,
const int32_t input_dims,
const int32_t *input_shape,
const int32_t axis,
const int32_t num_splits,
const int32_t *split_dims,
int16_t *const *output_data
)
```
int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis
:::note
This function, data layout independent, can be used to split either int8 or uint8 tensors because it does not involve any arithmetic operation.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to the flattened input tensor data. |
| input_dims | const int32_t | in | Number of dimensions in input_shape. |
| input_shape | const int32_t * | in | Array of length input_dims describing the shape of input_data. |
| axis | const int32_t | in | Axis along which to split (0 <= axis < input_dims). |
| num_splits | const int32_t | in | Number of output tensors to produce. |
| split_dims | const int32_t * | in | Array of length num_splits giving size of each slice along axis. |
| output_data | int16_t *const * | out | Array of pointers; output_data[i] points to storage for the i-th output tensor. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR if split_dims sum mismatch. |
Source: `Include/arm_nnfunctions.h:6570`
## arm_svdf_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_svdf_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *input_ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_svdf_params *svdf_params,
const cmsis_nn_per_tensor_quant_params *input_quant_params,
const cmsis_nn_per_tensor_quant_params *output_quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *state_dims,
int8_t *state_data,
const cmsis_nn_dims *weights_feature_dims,
const int8_t *weights_feature_data,
const cmsis_nn_dims *weights_time_dims,
const int8_t *weights_time_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 SVDF function with 8 bit state tensor and 8 bit time weights
1. Supported framework: TensorFlow Lite micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Precomputed per-feature-batch kernel sums, supplied by the caller. This is an input the function only reads, not scratch it fills: an allocated but unfilled buffer yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. Mandatory on builds with the MVE extension (ARM_MATH_MVEI), where a NULL ctx->buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. Unused on every other build, where ctx->buf may be NULL. Sized by arm_svdf_s8_get_buffer_size(weights_feature_dims): weights_feature_dims->n * sizeof(int32_t) where the sums are used, 0 otherwise. Note this is weights_feature_dims->n, not a filter_dims->c - do not size this buffer with `arm_fully_connected_s8_get_buffer_size()`, which reads a different field and under-allocates. Fill it with arm_vector_sum_s8(ctx->buf, input_dims->h, weights_feature_dims->n, weights_feature_data, -svdf_params->input_offset, 0, NULL) so that entry j holds -input_offset * sum(weights_feature row j). The contents depend only on weights_feature_data and svdf_params->input_offset, so they may be computed once at load time and reused across calls until one of those changes. The buffer is specific to one layer's weights and cannot be shared between layers. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call. |
| input_ctx | const cmsis_nn_context * | in | Scratch buffer written by this function, holding one int32_t accumulator per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_s8_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(int32_t) bytes, the same figure on every build target. This function does not read input_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons. |
| output_ctx | const cmsis_nn_context * | in | Scratch buffer written by this function, holding one int32_t accumulator per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_s8_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t) bytes, truncating division, the same figure on every build target. This function does not read output_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons. |
| svdf_params | const cmsis_nn_svdf_params * | in | SVDF Parameters Range of svdf_params->input_offset : [-128, 127] Range of svdf_params->output_offset : [-128, 127] |
| input_quant_params | const cmsis_nn_per_tensor_quant_params * | in | Input quantization parameters |
| output_quant_params | const cmsis_nn_per_tensor_quant_params * | in | Output quantization parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| input_data | const int8_t * | in | Pointer to input tensor |
| state_dims | const cmsis_nn_dims * | in | State tensor dimensions |
| state_data | int8_t * | in, out | Pointer to state tensor |
| weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions |
| weights_feature_data | const int8_t * | in | Pointer to the weights (feature) tensor |
| weights_time_dims | const cmsis_nn_dims * | in | Weights (time) tensor dimensions |
| weights_time_data | const int8_t * | in | Pointer to the weights (time) tensor |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions |
| bias_data | const int32_t * | in | Pointer to bias tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| output_data | int8_t * | out | Pointer to the output tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Source: `Include/arm_nnfunctions.h:6662`
## arm_svdf_state_s16_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_svdf_state_s16_s8(
const cmsis_nn_context *input_ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_svdf_params *svdf_params,
const cmsis_nn_per_tensor_quant_params *input_quant_params,
const cmsis_nn_per_tensor_quant_params *output_quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *state_dims,
int16_t *state_data,
const cmsis_nn_dims *weights_feature_dims,
const int8_t *weights_feature_data,
const cmsis_nn_dims *weights_time_dims,
const int16_t *weights_time_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)
```
s8 SVDF function with 16 bit state tensor and 16 bit time weights
1. Supported framework: TensorFlow Lite micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_ctx | const cmsis_nn_context * | in | Scratch buffer written by this function, holding one int32_t accumulator per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_state_s16_s8_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(int32_t) bytes, the same figure on every build target. Note the accumulators are int32_t even though the state tensor is int16_t - this buffer does not shrink with the state width. This function does not read input_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons. |
| output_ctx | const cmsis_nn_context * | in | Scratch buffer written by this function, holding one int32_t accumulator per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_state_s16_s8_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t) bytes, truncating division, the same figure on every build target. This function does not read output_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons. |
| svdf_params | const cmsis_nn_svdf_params * | in | SVDF Parameters Range of svdf_params->input_offset : [-128, 127] Range of svdf_params->output_offset : [-128, 127] |
| input_quant_params | const cmsis_nn_per_tensor_quant_params * | in | Input quantization parameters |
| output_quant_params | const cmsis_nn_per_tensor_quant_params * | in | Output quantization parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| input_data | const int8_t * | in | Pointer to input tensor |
| state_dims | const cmsis_nn_dims * | in | State tensor dimensions |
| state_data | int16_t * | in, out | Pointer to state tensor |
| weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions |
| weights_feature_data | const int8_t * | in | Pointer to the weights (feature) tensor |
| weights_time_dims | const cmsis_nn_dims * | in | Weights (time) tensor dimensions |
| weights_time_data | const int16_t * | in | Pointer to the weights (time) tensor |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions |
| bias_data | const int32_t * | in | Pointer to bias tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| output_data | int8_t * | out | Pointer to the output tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:6734`
## arm_svdf_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_svdf_s8_get_buffer_size(const cmsis_nn_dims *weights_feature_dims)
```
Get size of the kernel-sum buffer required by `arm_svdf_s8()`.
For a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. `arm_svdf_s8()` has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other `cmsis_nn_dims` of that call can size it - in particular `arm_fully_connected_s8_get_buffer_size()` reads a different field and under-allocates. See `arm_svdf_s8()` for the buffer's layout, how to fill it and when it may be reused.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| weights_feature_dims | const cmsis_nn_dims * | in | dimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:6767`
## arm_svdf_s8_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_svdf_s8_get_buffer_size_dsp(const cmsis_nn_dims *weights_feature_dims)
```
Get size of the kernel-sum buffer required by `arm_svdf_s8()` for processors with DSP extension.
For a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. `arm_svdf_s8()` has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other `cmsis_nn_dims` of that call can size it - in particular `arm_fully_connected_s8_get_buffer_size()` reads a different field and under-allocates. See `arm_svdf_s8()` for the buffer's layout, how to fill it and when it may be reused.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_svdf_s8_get_buffer_size()`.
:::
:::note
Also validates dims like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| weights_feature_dims | const cmsis_nn_dims * | in | dimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:6778`
## arm_svdf_s8_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_svdf_s8_get_buffer_size_mve(const cmsis_nn_dims *weights_feature_dims)
```
Get size of the kernel-sum buffer required by `arm_svdf_s8()` for Arm(R) Helium Architecture case.
For a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. `arm_svdf_s8()` has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other `cmsis_nn_dims` of that call can size it - in particular `arm_fully_connected_s8_get_buffer_size()` reads a different field and under-allocates. See `arm_svdf_s8()` for the buffer's layout, how to fill it and when it may be reused.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_svdf_s8_get_buffer_size()`.
:::
:::note
Also validates dims like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| weights_feature_dims | const cmsis_nn_dims * | in | dimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:6789`
## arm_svdf_s8_input_ctx_get_buffer_size
`function` · `c`
```c
int32_t arm_svdf_s8_input_ctx_get_buffer_size(
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)
```
Get size of the input_ctx staging buffer required by `arm_svdf_s8()`.
Returns input_dims->n * weights_feature_dims->n * sizeof(int32_t). Unlike `arm_svdf_s8_get_buffer_size()`, this figure does not vary by build target: `arm_svdf_s8()` stages this buffer on every build, not only under MVE, so there is no _dsp / _mve pair to choose between and the validation runs on every target.
:::note
This is a different buffer from the one `arm_svdf_s8_get_buffer_size()` describes. That one sizes the read-only kernel sums passed as ctx; this one sizes the scratch passed as input_ctx.
:::
:::note
0 is a valid return for a degenerate shape (input_dims->n == 0). Unlike the general rule in README.md, a 0 here does NOT mean you may pass { NULL, 0 }: `arm_svdf_s8()` rejects a NULL input_ctx->buf with ARM_CMSIS_NN_ARG_ERROR regardless of the size. -1 is used only for an out-of-range or NULL argument.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
| weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:6812`
## arm_svdf_s8_output_ctx_get_buffer_size
`function` · `c`
```c
int32_t arm_svdf_s8_output_ctx_get_buffer_size(
const cmsis_nn_svdf_params *svdf_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)
```
Get size of the output_ctx staging buffer required by `arm_svdf_s8()`.
Returns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t). The division truncates, matching the kernel's own unit count. As with `arm_svdf_s8_input_ctx_get_buffer_size()`, the figure is the same on every build target and the validation runs on every target.
:::note
Same degenerate-0 contract as `arm_svdf_s8_input_ctx_get_buffer_size()`, including that a 0 does not license passing { NULL, 0 }. A rank greater than weights_feature_dims->n truncates the unit count to 0 and so returns 0.
:::
:::note
`arm_svdf_s8()` narrows svdf_params->rank to int16_t before dividing by it, so a rank outside int16_t range would make this query and the kernel disagree - 65538 narrows to 2. The kernel can then write unboundedly more than the untruncated formula reports, because that formula truncates to 0 whenever weights_feature_dims->n < 65538: at weights_feature_dims->n = 100 it would report 0 bytes while the kernel writes 50 units, i.e. 200 bytes. Such a rank returns -1 rather than a number the kernel will not honour. Ranks that survive the int16_t round trip, that is within [-32768, 32767], are unaffected; this library does not otherwise constrain svdf_params->rank.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| svdf_params | const cmsis_nn_svdf_params * | in | SVDF parameters; only svdf_params->rank is read |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
| weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any pointer is NULL, if svdf_params->rank is zero, negative or outside int16_t range, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:6842`
## arm_svdf_state_s16_s8_input_ctx_get_buffer_size
`function` · `c`
```c
int32_t arm_svdf_state_s16_s8_input_ctx_get_buffer_size(
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)
```
Get size of the input_ctx staging buffer required by `arm_svdf_state_s16_s8()`.
Returns input_dims->n * weights_feature_dims->n * sizeof(int32_t). Unlike `arm_svdf_s8_get_buffer_size()`, this figure does not vary by build target: `arm_svdf_s8()` stages this buffer on every build, not only under MVE, so there is no _dsp / _mve pair to choose between and the validation runs on every target.
:::note
This is a different buffer from the one `arm_svdf_s8_get_buffer_size()` describes. That one sizes the read-only kernel sums passed as ctx; this one sizes the scratch passed as input_ctx.
:::
:::note
0 is a valid return for a degenerate shape (input_dims->n == 0). Unlike the general rule in README.md, a 0 here does NOT mean you may pass { NULL, 0 }: `arm_svdf_s8()` rejects a NULL input_ctx->buf with ARM_CMSIS_NN_ARG_ERROR regardless of the size. -1 is used only for an out-of-range or NULL argument.
:::
Returns input_dims->n * weights_feature_dims->n * sizeof(int32_t) - the same figure as `arm_svdf_s8_input_ctx_get_buffer_size()` for the same shape. The accumulators are int32_t even though `arm_svdf_state_s16_s8()` carries an int16_t state tensor, so this buffer does not shrink with the state width.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
| weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:6855`
## arm_svdf_state_s16_s8_output_ctx_get_buffer_size
`function` · `c`
```c
int32_t arm_svdf_state_s16_s8_output_ctx_get_buffer_size(
const cmsis_nn_svdf_params *svdf_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)
```
Get size of the output_ctx staging buffer required by `arm_svdf_state_s16_s8()`.
Returns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t). The division truncates, matching the kernel's own unit count. As with `arm_svdf_s8_input_ctx_get_buffer_size()`, the figure is the same on every build target and the validation runs on every target.
:::note
Same degenerate-0 contract as `arm_svdf_s8_input_ctx_get_buffer_size()`, including that a 0 does not license passing { NULL, 0 }. A rank greater than weights_feature_dims->n truncates the unit count to 0 and so returns 0.
:::
:::note
`arm_svdf_s8()` narrows svdf_params->rank to int16_t before dividing by it, so a rank outside int16_t range would make this query and the kernel disagree - 65538 narrows to 2. The kernel can then write unboundedly more than the untruncated formula reports, because that formula truncates to 0 whenever weights_feature_dims->n < 65538: at weights_feature_dims->n = 100 it would report 0 bytes while the kernel writes 50 units, i.e. 200 bytes. Such a rank returns -1 rather than a number the kernel will not honour. Ranks that survive the int16_t round trip, that is within [-32768, 32767], are unaffected; this library does not otherwise constrain svdf_params->rank.
:::
Returns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t), truncating division - the same figure as `arm_svdf_s8_output_ctx_get_buffer_size()` for the same shape.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| svdf_params | const cmsis_nn_svdf_params * | in | SVDF parameters; only svdf_params->rank is read |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
| weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any pointer is NULL, if svdf_params->rank is zero, negative or outside int16_t range, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:6865`
## arm_lstm_unidirectional_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_lstm_unidirectional_s8(
const int8_t *input,
int8_t *output,
const cmsis_nn_lstm_params *params,
cmsis_nn_lstm_context *buffers
)
```
LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input data |
| output | int8_t * | out | Pointer to output data |
| params | const cmsis_nn_lstm_params * | in | Struct containing all information about the lstm operator, see arm_nn_types. |
| buffers | cmsis_nn_lstm_context * | in, out | Struct containing pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. Size temp1 with `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` and temp2 with `arm_lstm_unidirectional_s8_temp2_get_buffer_size()` - both hold int16_t gate vectors even though the layer datatype is s8, so sizing them in s8 elements under-allocates by half. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:6892`
## arm_lstm_unidirectional_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_lstm_unidirectional_s16(
const int16_t *input,
int16_t *output,
const cmsis_nn_lstm_params *params,
cmsis_nn_lstm_context *buffers
)
```
LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to input data |
| output | int16_t * | out | Pointer to output data |
| params | const cmsis_nn_lstm_params * | in | Struct containing all information about the lstm operator, see arm_nn_types. |
| buffers | cmsis_nn_lstm_context * | in, out | Struct containing pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. Size temp1 with `arm_lstm_unidirectional_s16_temp1_get_buffer_size()` and temp2 with `arm_lstm_unidirectional_s16_temp2_get_buffer_size()`. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:6914`
## arm_lstm_unidirectional_s8_temp1_get_buffer_size
`function` · `c`
```c
int32_t arm_lstm_unidirectional_s8_temp1_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)
```
Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_s8()`.
:::note
time_steps does not enter the requirement: the buffer is reused by every step. A layer with time_steps == 0 runs no step and never dereferences the buffer, but the query still reports the per-step figure rather than 0.
:::
:::note
0 is only returned for a degenerate shape (batch_size or hidden_size of 0), for which `arm_lstm_unidirectional_s8()` makes no scratch access, so { NULL } is acceptable there per the README.md buffer convention. There is no runtime enforcement: the kernel does not range-check the buffers, and an undersized allocation is written past on every build target.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target. |
Source: `Include/arm_nnfunctions.h:6940`
## arm_lstm_unidirectional_s8_temp2_get_buffer_size
`function` · `c`
```c
int32_t arm_lstm_unidirectional_s8_temp2_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)
```
Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_s8()`.
:::note
time_steps does not enter the requirement: the buffer is reused by every step. A layer with time_steps == 0 runs no step and never dereferences the buffer, but the query still reports the per-step figure rather than 0.
:::
:::note
0 is only returned for a degenerate shape (batch_size or hidden_size of 0), for which `arm_lstm_unidirectional_s8()` makes no scratch access, so { NULL } is acceptable there per the README.md buffer convention. There is no runtime enforcement: the kernel does not range-check the buffers, and an undersized allocation is written past on every build target.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target. |
| | | Required buffer size in bytes: the same figure as `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` for the same params. temp2 stages the cell-gate vector and the tanh(cell_state) vector, both of the same extent as the gate vectors staged in temp1. |
Source: `Include/arm_nnfunctions.h:6950`
## arm_lstm_unidirectional_s16_temp1_get_buffer_size
`function` · `c`
```c
int32_t arm_lstm_unidirectional_s16_temp1_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)
```
Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_s16()`.
:::note
time_steps does not enter the requirement: the buffer is reused by every step. A layer with time_steps == 0 runs no step and never dereferences the buffer, but the query still reports the per-step figure rather than 0.
:::
:::note
0 is only returned for a degenerate shape (batch_size or hidden_size of 0), for which `arm_lstm_unidirectional_s8()` makes no scratch access, so { NULL } is acceptable there per the README.md buffer convention. There is no runtime enforcement: the kernel does not range-check the buffers, and an undersized allocation is written past on every build target.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target. |
| | | Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t) - the same figure as `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` for the same params, since both layer datatypes stage int16_t gate vectors. |
Source: `Include/arm_nnfunctions.h:6961`
## arm_lstm_unidirectional_s16_temp2_get_buffer_size
`function` · `c`
```c
int32_t arm_lstm_unidirectional_s16_temp2_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)
```
Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_s16()`.
:::note
time_steps does not enter the requirement: the buffer is reused by every step. A layer with time_steps == 0 runs no step and never dereferences the buffer, but the query still reports the per-step figure rather than 0.
:::
:::note
0 is only returned for a degenerate shape (batch_size or hidden_size of 0), for which `arm_lstm_unidirectional_s8()` makes no scratch access, so { NULL } is acceptable there per the README.md buffer convention. There is no runtime enforcement: the kernel does not range-check the buffers, and an undersized allocation is written past on every build target.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target. |
| | | Required buffer size in bytes: the same figure as `arm_lstm_unidirectional_s16_temp1_get_buffer_size()` for the same params. |
Source: `Include/arm_nnfunctions.h:6970`
## arm_batch_matmul_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_batch_matmul_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_bmm_params *bmm_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_lhs_dims,
const int8_t *input_lhs,
const cmsis_nn_dims *input_rhs_dims,
const int8_t *input_rhs,
const cmsis_nn_dims *output_dims,
int8_t *output
)
```
Batch matmul function with 8 bit input and output.
1. Supported framework: TensorFlow Lite Micro
2. Performs row * row matrix multiplication with the RHS transposed.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Temporary scratch buffer for the per-row kernel sums of the RHS. Mandatory on builds with the MVE extension (ARM_MATH_MVEI), where a NULL ctx->buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. Unused on every other build, where ctx->buf may be NULL. Sized by arm_batch_matmul_s8_get_buffer_size(input_rhs_dims) - pass the same input_rhs_dims given below. That is input_rhs_dims->w * sizeof(int32_t) where the sums are used, 0 otherwise. Do not size this buffer with `arm_fully_connected_s8_get_buffer_size()`: it reads a different field, and an allocation short of input_rhs_dims->w words is written past its end. The function fills the buffer itself before each use, so the caller does not need to initialize it. ctx->buf must be aligned to sizeof(int32_t). If ctx->size is non-zero it is validated against the requirement and a buffer too small is rejected with ARM_CMSIS_NN_ARG_ERROR; a ctx->size of 0 skips that check. A negative input_rhs_dims->w, or one large enough that the required size exceeds INT32_MAX, is rejected with ARM_CMSIS_NN_ARG_ERROR regardless of ctx->size. The caller is expected to clear the buffer, if applicable, for security reasons. |
| bmm_params | const cmsis_nn_bmm_params * | in | Batch matmul Parameters Adjoint flags are currently unused and do not transpose either input; callers must supply the tensors in the layouts described below. |
| quant_params | const cmsis_nn_per_tensor_quant_params * | in | Quantization parameters |
| input_lhs_dims | const cmsis_nn_dims * | in | Input lhs tensor dimensions. This s8 function treats w as the row count and c as the inner dimension. This differs from `arm_batch_matmul_f32()`, so its dimension mapping must not be reused here. |
| input_lhs | const int8_t * | in | Pointer to input tensor |
| input_rhs_dims | const cmsis_nn_dims * | in | Input rhs tensor dimensions. The RHS must already be transposed, with w as its row count and c equal to input_lhs_dims->c. |
| input_rhs | const int8_t * | in | Pointer to transposed input tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| output | int8_t * | out | Pointer to the output tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the following:
- `ARM_CMSIS_NN_ARG_ERROR` if an MVE build receives an invalid context, a negative or unrepresentable RHS row count, or a declared context size below the requirement. - `ARM_CMSIS_NN_SUCCESS` on success. |
Source: `Include/arm_nnfunctions.h:7017`
## arm_batch_matmul_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_batch_matmul_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_bmm_params *bmm_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_lhs_dims,
const int16_t *input_lhs,
const cmsis_nn_dims *input_rhs_dims,
const int16_t *input_rhs,
const cmsis_nn_dims *output_dims,
int16_t *output
)
```
Batch matmul function with 16 bit input and output.
1. Supported framework: TensorFlow Lite Micro
2. Performs row * row matrix multiplication with the RHS transposed.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Unused: this function requires no scratch buffer and does not read or write ctx on any build, so ctx->buf may be NULL. Retained for signature compatibility with `arm_batch_matmul_s8()`. There is deliberately no arm_batch_matmul_s16_get_buffer_size(); in particular `arm_fully_connected_s8_get_buffer_size()` is not the sizer for this argument. If a real buffer is passed, the caller is expected to clear it, if applicable, for security reasons. |
| bmm_params | const cmsis_nn_bmm_params * | in | Batch matmul Parameters Adjoint flags are currently unused. |
| quant_params | const cmsis_nn_per_tensor_quant_params * | in | Quantization parameters |
| input_lhs_dims | const cmsis_nn_dims * | in | Input lhs tensor dimensions. This should be NHWC where LHS.C = RHS.C |
| input_lhs | const int16_t * | in | Pointer to input tensor |
| input_rhs_dims | const cmsis_nn_dims * | in | Input lhs tensor dimensions. This is expected to be transposed so should be NHWC where LHS.C = RHS.C |
| input_rhs | const int16_t * | in | Pointer to transposed input tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| output | int16_t * | out | Pointer to the output tensor |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7057`
## arm_batch_matmul_s8_get_buffer_size
`function` · `c`
```c
int32_t arm_batch_matmul_s8_get_buffer_size(const cmsis_nn_dims *input_rhs_dims)
```
Get size of the scratch buffer required by `arm_batch_matmul_s8()`.
For a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with `arm_fully_connected_s8_get_buffer_size()`, which reads .c) writes past the allocation whenever the rhs has more rows than columns. `arm_batch_matmul_s16()` needs no scratch buffer and so has no corresponding sizer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_rhs_dims | const cmsis_nn_dims * | in | dimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:7082`
## arm_batch_matmul_s8_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_batch_matmul_s8_get_buffer_size_dsp(const cmsis_nn_dims *input_rhs_dims)
```
Get size of the scratch buffer required by `arm_batch_matmul_s8()` for processors with DSP extension.
For a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with `arm_fully_connected_s8_get_buffer_size()`, which reads .c) writes past the allocation whenever the rhs has more rows than columns. `arm_batch_matmul_s16()` needs no scratch buffer and so has no corresponding sizer.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_batch_matmul_s8_get_buffer_size()`.
:::
:::note
Also validates dims like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_rhs_dims | const cmsis_nn_dims * | in | dimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:7093`
## arm_batch_matmul_s8_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_batch_matmul_s8_get_buffer_size_mve(const cmsis_nn_dims *input_rhs_dims)
```
Get size of the scratch buffer required by `arm_batch_matmul_s8()` for Arm(R) Helium Architecture case.
For a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with `arm_fully_connected_s8_get_buffer_size()`, which reads .c) writes past the allocation whenever the rhs has more rows than columns. `arm_batch_matmul_s16()` needs no scratch buffer and so has no corresponding sizer.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_batch_matmul_s8_get_buffer_size()`.
:::
:::note
Also validates dims like the top-level dispatcher, returning -1 for invalid values.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_rhs_dims | const cmsis_nn_dims * | in | dimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()` |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnfunctions.h:7104`
## arm_pad_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_pad_s8(
const int8_t *input,
int8_t *output,
const int8_t pad_value,
const cmsis_nn_dims *input_size,
const cmsis_nn_dims *pre_pad,
const cmsis_nn_dims *post_pad
)
```
Expands the size of the input by adding constant values before and after the data, in all dimensions.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input data |
| output | int8_t * | out | Pointer to output data |
| pad_value | const int8_t | in | Value to pad with |
| input_size | const cmsis_nn_dims * | in | Input tensor dimensions |
| pre_pad | const cmsis_nn_dims * | in | Padding to apply before data in each dimension |
| post_pad | const cmsis_nn_dims * | in | Padding to apply after data in each dimension |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7124`
## arm_pad_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_pad_s16(
const int16_t *input,
int16_t *output,
const int16_t pad_value,
const cmsis_nn_dims *input_size,
const cmsis_nn_dims *pre_pad,
const cmsis_nn_dims *post_pad
)
```
Expands the size of the input by adding constant values before and after the data, in all dimensions.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to input data |
| output | int16_t * | out | Pointer to output data |
| pad_value | const int16_t | in | Value to pad with |
| input_size | const cmsis_nn_dims * | in | Input tensor dimensions |
| pre_pad | const cmsis_nn_dims * | in | Padding to apply before data in each dimension |
| post_pad | const cmsis_nn_dims * | in | Padding to apply after data in each dimension |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7144`
## arm_mean_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_mean_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t input_offset,
const cmsis_nn_dims *axis_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift
)
```
Computes the mean of the input tensor along the specified axis. The output multipler and shift must have the output count folded into them.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| input_offset | const int32_t | in | Input offset |
| axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
| output_data | int8_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| out_offset | const int32_t | in | Output offset |
| out_mult | const int32_t | in | Output quantization multiplier |
| out_shift | const int32_t | in | Output quantization shift |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7173`
## arm_mean_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_mean_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t input_offset,
const cmsis_nn_dims *axis_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift
)
```
Computes the mean of the input tensor along the specified axis. The output multipler and shift must have the output count folded into them.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| input_offset | const int32_t | in | Input offset |
| axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
| output_data | int16_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
| out_offset | const int32_t | in | Output offset |
| out_mult | const int32_t | in | Output quantization multiplier |
| out_shift | const int32_t | in | Output quantization shift |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7200`
## arm_argmax_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_argmax_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t axis,
int32_t *output_data
)
```
Compute ArgMax indices of an s8 tensor along a specific axis.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (NHWC layout) |
| axis | const int32_t | in | Reduction axis in range [0, 3] |
| output_data | int32_t * | out | Pointer to output indices (int32_t) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`. |
Source: `Include/arm_nnfunctions.h:7221`
## arm_argmin_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_argmin_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t axis,
int32_t *output_data
)
```
Compute ArgMin indices of an s8 tensor along a specific axis.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (NHWC layout) |
| axis | const int32_t | in | Reduction axis in range [0, 3] |
| output_data | int32_t * | out | Pointer to output indices (int32_t) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`. |
Source: `Include/arm_nnfunctions.h:7234`
## arm_argmax_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_argmax_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t axis,
int32_t *output_data
)
```
Compute ArgMax indices of an s16 tensor along a specific axis.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (NHWC layout) |
| axis | const int32_t | in | Reduction axis in range [0, 3] |
| output_data | int32_t * | out | Pointer to output indices (int32_t) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`. |
Source: `Include/arm_nnfunctions.h:7247`
## arm_argmin_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_argmin_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t axis,
int32_t *output_data
)
```
Compute ArgMin indices of an s16 tensor along a specific axis.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (NHWC layout) |
| axis | const int32_t | in | Reduction axis in range [0, 3] |
| output_data | int32_t * | out | Pointer to output indices (int32_t) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`. |
Source: `Include/arm_nnfunctions.h:7260`
## arm_reduce_max_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_max_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Computes the max of the input tensor along the specified axis.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
| output_data | int8_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7274`
## arm_reduce_max_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_max_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Computes the max of the input tensor along the specified axis.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
| output_data | int16_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7292`
## arm_reduce_min_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_min_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Computes the min of the input tensor along the specified axis.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
| output_data | int8_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7310`
## arm_reduce_min_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_reduce_min_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Computes the min of the input tensor along the specified axis.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
| output_data | int16_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7328`
## arm_quantize_f32_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_quantize_f32_s8(
const float *input,
int8_t *output,
int32_t size,
int32_t zero_point,
float scale
)
```
Quantize a floating-point array into int8_t format.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float * | in | Pointer to the input float array. |
| output | int8_t * | out | Pointer to the output int8_t array. |
| size | int32_t | in | Number of elements in the arrays. |
| zero_point | int32_t | in | Zero point (offset) to apply during quantization. |
| scale | float | in | Scale factor to apply during quantization. Must be a positive finite number. A scale that is zero, negative, NaN, Inf, or small enough that its reciprocal overflows is unsupported, and the result is then unspecified: the scalar and Helium legs are not guaranteed to agree for such a scale. Denormal inputs are likewise unspecified: the Helium leg flushes them to zero, so the two legs are not guaranteed to agree for a denormal input at any scale. Screen denormals if you need them to match. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when `zero_point` lies outside the int8_t range. Values round half away from zero and saturate to the int8_t range after the zero point is applied; NaN maps to `zero_point`. |
Source: `Include/arm_nnfunctions.h:7356`
## arm_quantize_f32_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_quantize_f32_s16(
const float *input,
int16_t *output,
int32_t size,
int32_t zero_point,
float scale
)
```
Quantize a floating-point array into int16_t format.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float * | in | Pointer to the input float array. |
| output | int16_t * | out | Pointer to the output int16_t array. |
| size | int32_t | in | Number of elements in the arrays. |
| zero_point | int32_t | in | Zero point (offset) to apply during quantization. |
| scale | float | in | Scale factor to apply during quantization. Must be a positive finite number. A scale that is zero, negative, NaN, Inf, or small enough that its reciprocal overflows is unsupported, and the result is then unspecified: the scalar and Helium legs are not guaranteed to agree for such a scale. Denormal inputs are likewise unspecified: the Helium leg flushes them to zero, so the two legs are not guaranteed to agree for a denormal input at any scale. Screen denormals if you need them to match. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when `zero_point` lies outside the int16_t range. Values round half away from zero and saturate to the int16_t range after the zero point is applied; NaN maps to `zero_point`. |
Source: `Include/arm_nnfunctions.h:7376`
## arm_requantize_s8_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_requantize_s8_s8(
const int8_t *input,
int8_t *output,
int32_t size,
int32_t effective_scale_multiplier,
int32_t effective_scale_shift,
int32_t input_zeropoint,
int32_t output_zeropoint
)
```
Requantize an int8_t array to another int8_t range with a different scale.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input int8_t array. |
| output | int8_t * | out | Pointer to the output int8_t array. |
| size | int32_t | in | Number of elements in the arrays. |
| effective_scale_multiplier | int32_t | in | Multiplier used for the scaling operation. |
| effective_scale_shift | int32_t | in | Right or left shift (depending on sign) applied after the multiplier. |
| input_zeropoint | int32_t | in | Zero point of the input data. |
| output_zeropoint | int32_t | in | Zero point of the output data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7390`
## arm_requantize_s16_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_requantize_s16_s16(
const int16_t *input,
int16_t *output,
int32_t size,
int32_t effective_scale_multiplier,
int32_t effective_scale_shift,
int32_t input_zeropoint,
int32_t output_zeropoint
)
```
Requantize an int16_t array to another int16_t range with a different scale.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input int16_t array. |
| output | int16_t * | out | Pointer to the output int16_t array. |
| size | int32_t | in | Number of elements in the arrays. |
| effective_scale_multiplier | int32_t | in | Multiplier used for the scaling operation. |
| effective_scale_shift | int32_t | in | Right or left shift (depending on sign) applied after the multiplier. |
| input_zeropoint | int32_t | in | Zero point of the input data. |
| output_zeropoint | int32_t | in | Zero point of the output data. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7410`
## arm_dequantize_s8_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_dequantize_s8_f32(
const int8_t *input,
float *output,
int32_t size,
int32_t zero_point,
float scale
)
```
Dequantize an int8_t array back to floating-point format.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input int8_t array. |
| output | float * | out | Pointer to the output float array. |
| size | int32_t | in | Number of elements in the arrays. |
| zero_point | int32_t | in | Zero point (offset) that was used during quantization. |
| scale | float | in | Scale factor that was used during quantization. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7429`
## arm_dequantize_s16_f32
`function` · `c`
```c
arm_cmsis_nn_status arm_dequantize_s16_f32(
const int16_t *input,
float *output,
int32_t size,
int32_t zero_point,
float scale
)
```
Dequantize an int16_t array back to floating-point format.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to the input int16_t array. |
| output | float * | out | Pointer to the output float array. |
| size | int32_t | in | Number of elements in the arrays. |
| zero_point | int32_t | in | Zero point (offset) that was used during quantization. |
| scale | float | in | Scale factor that was used during quantization. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7442`
## arm_strided_slice_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_strided_slice_s8(
const int8_t *input_data,
int8_t *output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const begin_dims,
const cmsis_nn_dims *const stride_dims,
const cmsis_nn_dims *const output_dims
)
```
Strided slice function for int8 data.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to input tensor |
| output_data | int8_t * | out | Pointer to output tensor |
| input_dims | const cmsis_nn_dims *const | in | Input tensor dimensions |
| begin_dims | const cmsis_nn_dims *const | in | Begin dimensions for slicing |
| stride_dims | const cmsis_nn_dims *const | in | Stride dimensions for slicing |
| output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7465`
## arm_strided_slice_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_strided_slice_s16(
const int16_t *input_data,
int16_t *output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const begin_dims,
const cmsis_nn_dims *const stride_dims,
const cmsis_nn_dims *const output_dims
)
```
Strided slice function for int16 data.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to input tensor |
| output_data | int16_t * | out | Pointer to output tensor |
| input_dims | const cmsis_nn_dims *const | in | Input tensor dimensions |
| begin_dims | const cmsis_nn_dims *const | in | Begin dimensions for slicing |
| stride_dims | const cmsis_nn_dims *const | in | Stride dimensions for slicing |
| output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7488`
## arm_strided_slice_s32
`function` · `c`
```c
arm_cmsis_nn_status arm_strided_slice_s32(
const int32_t *input_data,
int32_t *output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const begin_dims,
const cmsis_nn_dims *const stride_dims,
const cmsis_nn_dims *const output_dims
)
```
Strided slice function for int32 data.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int32_t * | in | Pointer to input tensor |
| output_data | int32_t * | out | Pointer to output tensor |
| input_dims | const cmsis_nn_dims *const | in | Input tensor dimensions |
| begin_dims | const cmsis_nn_dims *const | in | Begin dimensions for slicing |
| stride_dims | const cmsis_nn_dims *const | in | Stride dimensions for slicing |
| output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7511`
## arm_gather_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_gather_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_params *params,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Gather elements along an axis for int8 tensors.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int8_t * | in | Pointer to input tensor data |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| indices_data | const int32_t * | in | Pointer to indices tensor data (int32) |
| indices_dims | const cmsis_nn_dims * | in | Indices tensor dimensions |
| params | const cmsis_nn_gather_params * | in | Pointer to gather parameters |
| output_data | int8_t * | out | Pointer to output tensor data |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7540`
## arm_gather_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_gather_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_params *params,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Gather elements along an axis for int16 tensors.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const int16_t * | in | Pointer to input tensor data |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
| indices_data | const int32_t * | in | Pointer to indices tensor data (int32) |
| indices_dims | const cmsis_nn_dims * | in | Indices tensor dimensions |
| params | const cmsis_nn_gather_params * | in | Pointer to gather parameters |
| output_data | int16_t * | out | Pointer to output tensor data |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7565`
## arm_gather_nd_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_gather_nd_s8(
const int8_t *params_data,
const cmsis_nn_dims *params_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_nd_params *params,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Gather_nd slices for int8 tensors.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| params_data | const int8_t * | in | Pointer to params tensor data |
| params_dims | const cmsis_nn_dims * | in | Params tensor dimensions |
| indices_data | const int32_t * | in | Pointer to indices tensor data (int32) |
| indices_dims | const cmsis_nn_dims * | in | Indices tensor dimensions |
| params | const cmsis_nn_gather_nd_params * | in | Pointer to gather_nd parameters |
| output_data | int8_t * | out | Pointer to output tensor data |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7590`
## arm_gather_nd_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_gather_nd_s16(
const int16_t *params_data,
const cmsis_nn_dims *params_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_nd_params *params,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)
```
Gather_nd slices for int16 tensors.
1. Supported framework: TensorFlow Lite Micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| params_data | const int16_t * | in | Pointer to params tensor data |
| params_dims | const cmsis_nn_dims * | in | Params tensor dimensions |
| indices_data | const int32_t * | in | Pointer to indices tensor data (int32) |
| indices_dims | const cmsis_nn_dims * | in | Indices tensor dimensions |
| params | const cmsis_nn_gather_nd_params * | in | Pointer to gather_nd parameters |
| output_data | int16_t * | out | Pointer to output tensor data |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7615`
## arm_tile_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_tile_s8(const int8_t *input, const cmsis_nn_tile_params *params, int8_t *output)
```
Tile an int8 tensor along each dimension.
1. Supported framework: TensorFlow Lite Micro
2. Maximum rank: 8
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input tensor data |
| params | const cmsis_nn_tile_params * | in | Pointer to tile parameters (rank, input_shape, multiples) |
| output | int8_t * | out | Pointer to output tensor data (pre-allocated by caller) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7642`
## arm_tile_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_tile_s16(const int16_t *input, const cmsis_nn_tile_params *params, int16_t *output)
```
Tile an int16 tensor along each dimension.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to input tensor data |
| params | const cmsis_nn_tile_params * | in | Pointer to tile parameters (rank, input_shape, multiples) |
| output | int16_t * | out | Pointer to output tensor data (pre-allocated by caller) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7654`
## arm_broadcast_to_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_broadcast_to_s8(const int8_t *input, const cmsis_nn_broadcast_to_params *params, int8_t *output)
```
Broadcast an int8 tensor to a target shape.
1. Input dimensions must be 1 or match the output dimension for each axis.
2. Maximum rank: 8
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input tensor data |
| params | const cmsis_nn_broadcast_to_params * | in | Pointer to broadcast parameters (rank, input/output shapes) |
| output | int8_t * | out | Pointer to output tensor data (pre-allocated by caller) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7676`
## arm_broadcast_to_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_broadcast_to_s16(
const int16_t *input,
const cmsis_nn_broadcast_to_params *params,
int16_t *output
)
```
Broadcast an int16 tensor to a target shape.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to input tensor data |
| params | const cmsis_nn_broadcast_to_params * | in | Pointer to broadcast parameters (rank, input/output shapes) |
| output | int16_t * | out | Pointer to output tensor data (pre-allocated by caller) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7689`
## arm_scatter_nd_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_scatter_nd_s8(
const int32_t *indices,
const int8_t *updates,
const cmsis_nn_scatter_nd_params *params,
int8_t *output
)
```
Scatter updates into a zero-initialized output tensor for int8.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| indices | const int32_t * | in | Pointer to indices data (int32, shape [num_updates, index_depth]) |
| updates | const int8_t * | in | Pointer to updates data |
| params | const cmsis_nn_scatter_nd_params * | in | Pointer to scatter_nd parameters |
| output | int8_t * | out | Pointer to output tensor data (pre-allocated by caller) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7707`
## arm_scatter_nd_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_scatter_nd_s16(
const int32_t *indices,
const int16_t *updates,
const cmsis_nn_scatter_nd_params *params,
int16_t *output
)
```
Scatter updates into a zero-initialized output tensor for int16.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| indices | const int32_t * | in | Pointer to indices data (int32, shape [num_updates, index_depth]) |
| updates | const int16_t * | in | Pointer to updates data |
| params | const cmsis_nn_scatter_nd_params * | in | Pointer to scatter_nd parameters |
| output | int16_t * | out | Pointer to output tensor data (pre-allocated by caller) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7723`
## arm_mirror_pad_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_mirror_pad_s8(const int8_t *input, const cmsis_nn_mirror_pad_params *params, int8_t *output)
```
Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input tensor data |
| params | const cmsis_nn_mirror_pad_params * | in | Pointer to mirror_pad parameters |
| output | int8_t * | out | Pointer to output tensor data (pre-allocated by caller) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7743`
## arm_mirror_pad_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_mirror_pad_s16(const int16_t *input, const cmsis_nn_mirror_pad_params *params, int16_t *output)
```
Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to input tensor data |
| params | const cmsis_nn_mirror_pad_params * | in | Pointer to mirror_pad parameters |
| output | int16_t * | out | Pointer to output tensor data (pre-allocated by caller) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7755`
## arm_where_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_where_s8(
const int8_t *condition,
const cmsis_nn_where_params *params,
int64_t *output,
int32_t *num_true
)
```
WHERE operator: return coordinates of non-zero elements in condition.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| condition | const int8_t * | in | Pointer to condition tensor data (int8, non-zero = true) |
| params | const cmsis_nn_where_params * | in | Pointer to where parameters (rank, shape) |
| output | int64_t * | out | Pointer to output coordinates (int64, shape [max_true, rank]) |
| num_true | int32_t * | out | Number of true elements found |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7774`
## arm_where_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_where_s16(
const int16_t *condition,
const cmsis_nn_where_params *params,
int64_t *output,
int32_t *num_true
)
```
WHERE operator: return coordinates of non-zero elements in condition (int16).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| condition | const int16_t * | in | Pointer to condition tensor data (int16, non-zero = true) |
| params | const cmsis_nn_where_params * | in | Pointer to where parameters (rank, shape) |
| output | int64_t * | out | Pointer to output coordinates (int64, shape [max_true, rank]) |
| num_true | int32_t * | out | Number of true elements found |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7788`
## arm_select_v2_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_select_v2_s8(
const bool *condition,
const int8_t *x,
const int8_t *y,
const cmsis_nn_select_v2_params *params,
int8_t *output
)
```
SELECT_V2 with broadcast for int8 tensors.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| condition | const bool * | in | Pointer to condition tensor data (bool) |
| x | const int8_t * | in | Pointer to x tensor data (selected when condition is true) |
| y | const int8_t * | in | Pointer to y tensor data (selected when condition is false) |
| params | const cmsis_nn_select_v2_params * | in | Pointer to select_v2 parameters (broadcast strides) |
| output | int8_t * | out | Pointer to output tensor data |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7802`
## arm_select_v2_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_select_v2_s16(
const bool *condition,
const int16_t *x,
const int16_t *y,
const cmsis_nn_select_v2_params *params,
int16_t *output
)
```
SELECT_V2 with broadcast for int16 tensors.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| condition | const bool * | in | Pointer to condition tensor data (bool) |
| x | const int16_t * | in | Pointer to x tensor data |
| y | const int16_t * | in | Pointer to y tensor data |
| params | const cmsis_nn_select_v2_params * | in | Pointer to select_v2 parameters (broadcast strides) |
| output | int16_t * | out | Pointer to output tensor data |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7820`
## arm_reverse_sequence_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_reverse_sequence_s8(
const int8_t *input,
const int32_t *seq_lengths,
const cmsis_nn_reverse_sequence_params *params,
int8_t *output
)
```
Reverse variable-length sequences along a dimension for int8.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to input tensor data |
| seq_lengths | const int32_t * | in | Pointer to per-batch sequence lengths (int32) |
| params | const cmsis_nn_reverse_sequence_params * | in | Pointer to reverse_sequence parameters |
| output | int8_t * | out | Pointer to output tensor data |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7842`
## arm_reverse_sequence_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_reverse_sequence_s16(
const int16_t *input,
const int32_t *seq_lengths,
const cmsis_nn_reverse_sequence_params *params,
int16_t *output
)
```
Reverse variable-length sequences along a dimension for int16.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int16_t * | in | Pointer to input tensor data |
| seq_lengths | const int32_t * | in | Pointer to per-batch sequence lengths (int32) |
| params | const cmsis_nn_reverse_sequence_params * | in | Pointer to reverse_sequence parameters |
| output | int16_t * | out | Pointer to output tensor data |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7858`
## arm_dynamic_update_slice_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_dynamic_update_slice_s8(
const int8_t *operand,
const int8_t *update,
const int32_t *start_indices,
const cmsis_nn_dynamic_update_slice_params *params,
int8_t *output
)
```
Update a slice of an int8 operand tensor at runtime-determined indices.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| operand | const int8_t * | in | Pointer to operand tensor data (copied to output first) |
| update | const int8_t * | in | Pointer to update tensor data |
| start_indices | const int32_t * | in | Pointer to start index per dimension (int32, length = rank) |
| params | const cmsis_nn_dynamic_update_slice_params * | in | Pointer to dynamic_update_slice parameters |
| output | int8_t * | out | Pointer to output tensor data |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7880`
## arm_dynamic_update_slice_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_dynamic_update_slice_s16(
const int16_t *operand,
const int16_t *update,
const int32_t *start_indices,
const cmsis_nn_dynamic_update_slice_params *params,
int16_t *output
)
```
Update a slice of an int16 operand tensor at runtime-determined indices.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| operand | const int16_t * | in | Pointer to operand tensor data (copied to output first) |
| update | const int16_t * | in | Pointer to update tensor data |
| start_indices | const int32_t * | in | Pointer to start index per dimension (int32, length = rank) |
| params | const cmsis_nn_dynamic_update_slice_params * | in | Pointer to dynamic_update_slice parameters |
| output | int16_t * | out | Pointer to output tensor data |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnfunctions.h:7898`
---
# heliaCORE.arm_nnsupportfunctions
## USE_FAST_DW_CONV_S16_FUNCTION
`macro` · `c`
```c
#define USE_FAST_DW_CONV_S16_FUNCTION(dw_conv_params, filter_dims, input_dims, output_dims) (dw_conv_params->ch_mult == 1 && \ arm_nn_dw_conv_opt_dilation_supported(dw_conv_params, input_dims, filter_dims, output_dims) && \ filter_dims->w * filter_dims->h < 512)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | | | |
| filter_dims | | | |
| input_dims | | | |
| output_dims | | | |
Source: `Include/arm_nnsupportfunctions.h:45`
## LEFT_SHIFT
`macro` · `c`
```c
#define LEFT_SHIFT(_shift) (_shift > 0 ? _shift : 0)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| _shift | | | |
Source: `Include/arm_nnsupportfunctions.h:50`
## RIGHT_SHIFT
`macro` · `c`
```c
#define RIGHT_SHIFT(_shift) (_shift > 0 ? 0 : -_shift)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| _shift | | | |
Source: `Include/arm_nnsupportfunctions.h:51`
## MASK_IF_ZERO
`macro` · `c`
```c
#define MASK_IF_ZERO(x) (x) == 0 ? ~0 : 0
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | | | |
Source: `Include/arm_nnsupportfunctions.h:52`
## MASK_IF_NON_ZERO
`macro` · `c`
```c
#define MASK_IF_NON_ZERO(x) (x) != 0 ? ~0 : 0
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | | | |
Source: `Include/arm_nnsupportfunctions.h:53`
## SELECT_USING_MASK
`macro` · `c`
```c
#define SELECT_USING_MASK(mask, a, b) ((mask) & (a)) ^ (~(mask) & (b))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| mask | | | |
| a | | | |
| b | | | |
Source: `Include/arm_nnsupportfunctions.h:54`
## ARM_NN_MAX
`macro` · `c`
```c
#define ARM_NN_MAX(A, B) ((A) > (B) ? (A) : (B))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| A | | | |
| B | | | |
Source: `Include/arm_nnsupportfunctions.h:57`
## ARM_NN_MIN
`macro` · `c`
```c
#define ARM_NN_MIN(A, B) ((A) < (B) ? (A) : (B))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| A | | | |
| B | | | |
Source: `Include/arm_nnsupportfunctions.h:58`
## ARM_NN_CLAMP
`macro` · `c`
```c
#define ARM_NN_CLAMP(x, h, l) ARM_NN_MAX(ARM_NN_MIN((x), (h)), (l))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | | | |
| h | | | |
| l | | | |
Source: `Include/arm_nnsupportfunctions.h:59`
## arm_nn_min_f16h
`function` · `c`
```c
static _Float16 arm_nn_min_f16h(_Float16 a, _Float16 b)
```
Minimum of two scalar f16 values.
With ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 minNum via VMINNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MIN, an ordered compare, so its NaN handling depends on operand order: a NaN `b` is returned, a NaN `a` is not. Do not rely on NaN suppression on non-MVE builds.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | _Float16 | in | First operand |
| b | _Float16 | in | Second operand |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The smaller of `a` and `b` |
Source: `Include/arm_nnsupportfunctions.h:98`
## arm_nn_max_f16h
`function` · `c`
```c
static _Float16 arm_nn_max_f16h(_Float16 a, _Float16 b)
```
Maximum of two scalar f16 values.
With ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 maxNum via VMAXNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MAX, an ordered compare, so its NaN handling depends on operand order: a NaN `b` is returned, a NaN `a` is not. Do not rely on NaN suppression on non-MVE builds.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | _Float16 | in | First operand |
| b | _Float16 | in | Second operand |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The larger of `a` and `b` |
Source: `Include/arm_nnsupportfunctions.h:121`
## arm_nn_propagate_nan_f16h
`function` · `c`
```c
static _Float16 arm_nn_propagate_nan_f16h(_Float16 x, _Float16 y)
```
Returns `x` when `x` is NaN, otherwise `y`.
Both the NaN test and the select are performed on the bit patterns: the test is (bits & 0x7FFF) > 0x7C00 (all-ones exponent, non-zero mantissa), which is integer arithmetic that -ffinite-math-only (implied by the shipped -Ofast) has no license to fold, unlike the former floating-point self-compare `x != x` (#333 / #334); the bit-pattern select neither expands to an HFmode conditional move (PR target/118460) nor quiets/retags the NaN payload. This helper backs the f16 elementwise clamp and, via arm_nn_clamp_scalar_f16 / arm_nn_clamp_propagate_nan_f16h, the other f16 scalar clamp users arm_svdf_f16, arm_max_pool_f16 / arm_avg_pool_f16, the packed f16 matmul (arm_nn_mat_mult_nt_n_packed_f16), the scalar f16 RELU/RELU6/LEAKY_RELU activation legs, and arm_nn_vector_clamp_f16's scalar leg (conv/depthwise/transpose-conv f16, the 3x3 depthwise, and arm_nn_maxpool1d_f16) so wherever that scalar clamp runs, a NaN passes through it at every optimization level on the gated toolchains. That is a guarantee about the clamp, not the whole kernel: which builds run the scalar clamp, and whether a NaN survives the rest of the kernel to reach it, is per kernel several of these callers clamp with vmaxnmq/vminnmq on MVE builds (a NaN resolves to a bound there), and arm_max_pool_f16's max reduction drops a NaN before the clamp. The kernels with a NaN
:::note
(svdf, max/avg pool, packed matmul, activation) state their exact scope there; the arm_nn_vector_clamp_f16 family is covered by a test assertion in the transpose-conv f16 suite rather than per-kernel notes. The cortex-m55 MVE RELU/RELU6 f16 legs reach the same guarantee by the vector form of this idiom rather than by calling this helper: they restore NaN lanes with arm_nn_max_propagate_nan_mve_f16 / arm_nn_clamp_propagate_nan_mve_f16 (#382). The same idiom (bit-classified select) appears in arm_prelu_f16, which does not call this helper.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | _Float16 | in | Value whose NaN-ness selects the result. Returned unchanged when it is NaN. |
| y | _Float16 | in | Value returned when `x` is not NaN |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `x` if `x` is NaN, otherwise `y` |
Source: `Include/arm_nnsupportfunctions.h:167`
## arm_nn_clamp_f16h
`function` · `c`
```c
static _Float16 arm_nn_clamp_f16h(_Float16 x, _Float16 h, _Float16 l)
```
Drop-in equivalent of `ARM_NN_CLAMP(x, h, l)` for scalar _Float16 operands.
Includes the macro's NaN behaviour: `ARM_NN_MIN(NaN, h)` is h, so a NaN input resolves to the high bound, exactly as the macro does. Use arm_nn_clamp_propagate_nan_f16h() where TFLite NaN propagation is required.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | _Float16 | in | Value to clamp |
| h | _Float16 | in | Upper bound |
| l | _Float16 | in | Lower bound |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `x` clamped to [`l`, `h`] |
Source: `Include/arm_nnsupportfunctions.h:193`
## arm_nn_clamp_propagate_nan_f16h
`function` · `c`
```c
static _Float16 arm_nn_clamp_propagate_nan_f16h(_Float16 x, _Float16 l, _Float16 h)
```
Scalar f16 clamp with TFLite NaN semantics: NaN passes through unchanged.
Mirrors the MVE idiom in arm_nn_clamp_propagate_nan_mve_f16() (lower bound first, then upper bound, then restore NaN lanes). The NaN restore in arm_nn_propagate_nan_f16h() tests the integer bit pattern, so it holds at every optimization level including the shipped -Ofast; see #333 / #334. Bounds are assumed ordered (l <= h); inverted bounds are unspecified.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | _Float16 | in | Value to clamp |
| l | _Float16 | in | Lower bound |
| h | _Float16 | in | Upper bound |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `x` clamped to [`l`, `h`], or `x` itself when it is NaN |
Source: `Include/arm_nnsupportfunctions.h:212`
## arm_nn_abs_f16h
`function` · `c`
```c
static _Float16 arm_nn_abs_f16h(_Float16 x)
```
Absolute value of a scalar f16 value.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | _Float16 | in | Input value |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | \|`x`\| |
Source: `Include/arm_nnsupportfunctions.h:224`
## ARM_NN_ROUND_UP
`macro` · `c`
```c
#define ARM_NN_ROUND_UP(x, multiple) ((((x) + (multiple) - 1) / (multiple)) * (multiple))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | | | |
| multiple | | | |
Source: `Include/arm_nnsupportfunctions.h:236`
## REDUCE_MULTIPLIER
`macro` · `c`
```c
#define REDUCE_MULTIPLIER(_mult) ((_mult < 0x7FFF0000) ? ((_mult + (1 << 15)) >> 16) : 0x7FFF)
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| _mult | | | |
Source: `Include/arm_nnsupportfunctions.h:237`
## CH_IN_BLOCK_MVE
`macro` · `c`
```c
#define CH_IN_BLOCK_MVE (124)
```
Source: `Include/arm_nnsupportfunctions.h:245`
## S4_CH_IN_BLOCK_MVE
`macro` · `c`
```c
#define S4_CH_IN_BLOCK_MVE (124)
```
Source: `Include/arm_nnsupportfunctions.h:250`
## MAX_COL_COUNT
`macro` · `c`
```c
#define MAX_COL_COUNT (512)
```
Source: `Include/arm_nnsupportfunctions.h:254`
## REVERSE_TCOL_EFFICIENT_THRESHOLD
`macro` · `c`
```c
#define REVERSE_TCOL_EFFICIENT_THRESHOLD (16)
```
Source: `Include/arm_nnsupportfunctions.h:258`
## CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD
`macro` · `c`
```c
#define CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (1)
```
Source: `Include/arm_nnsupportfunctions.h:266`
## OPTIONAL_RESTRICT_KEYWORD
`macro` · `c`
```c
#define OPTIONAL_RESTRICT_KEYWORD
```
Source: `Include/arm_nnsupportfunctions.h:272`
## arm_nn_size_mul
`function` · `c`
```c
static int64_t arm_nn_size_mul(const int64_t acc, const int64_t factor)
```
Fold one dimension into a running buffer-size product, reporting overflow as -1.
Buffer-size queries return an int32_t byte count, so the product of the dimensions they multiply has to be rejected as soon as it cannot fit. Folding one factor at a time keeps the accumulator bounded: an accumulator already known to be <= INT32_MAX times a factor <= INT32_MAX cannot exceed about 2^62, so the int64_t accumulator itself never wraps. Chaining raw (int64_t) casts across three or more int32_t dims does not have that property - 65536 * 65536 * 65536 * 65536 is exactly 2^64 and folds back to 0, which would sail through a trailing "> INT32_MAX" test.
:::note
This is the -1 sentinel family, used by the s8/s16 integer buffer-size queries, by the eight SVDF staging queries (arm_svdf_{s8,state_s16_s8,f32,f16}_{input,output}_ctx_get_buffer_size) and by the s8/s16 LSTM temp-buffer queries and the GRU temp queries (arm_lstm_unidirectional_{s8,s16}_temp{1,2}_get_buffer_size, arm_gru_unidirectional_{f32,f16}_temp1_get_buffer_size). The four f32/f16 LSTM temp queries have no dimensions to fold (the buffers are unused) and answer -1 only for NULL params, 0 otherwise. It is not interchangeable with the arm_nn_checked_size_mul() / arm_nn_size_to_i32_or_zero() helpers in Source/NNSupportFunctions (shared header for the float sizers), which most f32 and f16 buffer-size queries use and which report an out-of-range size as 0. Mixing the two silently flips a sizer's out-of-range contract from "must never be used to size a buffer" to "you may pass { NULL, 0 }", so pick the one the surrounding family already uses.
:::
:::note
The split is per sizer, not per datatype. The four SVDF f32/f16 staging queries deliberately use this -1 family rather than the 0 one their neighbours use, because their kernels read ctx->size and a size of 0 opts out of the scratch-size check - so a 0-on-overflow answer fed back as { alloc(0), 0 } would disable the very check meant to catch it. Do not infer a sizer's sentinel from its datatype suffix.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| acc | const int64_t | in | Running product, or -1 if an earlier fold already overflowed. |
| factor | const int64_t | in | Next factor to fold in. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | acc * factor, or -1 if acc is already -1, factor is negative or out of int32_t range, or the product exceeds INT32_MAX. |
Source: `Include/arm_nnsupportfunctions.h:307`
## arm_nn_size_add
`function` · `c`
```c
static int64_t arm_nn_size_add(const int64_t acc, const int64_t addend)
```
Add to a running buffer-size product, reporting overflow as -1.
Companion to arm_nn_size_mul() for the sizers that append a fixed slack term.
:::note
Same sentinel caveat as arm_nn_size_mul() - see its note for the full split, including why the four SVDF f32/f16 staging queries use this -1 family rather than the 0-returning arm_nn_checked_size_mul() / arm_nn_size_to_i32_or_zero() family that most other float sizers use.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| acc | const int64_t | in | Running product, or -1 if an earlier step already overflowed. |
| addend | const int64_t | in | Value to add. Must be non-negative. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | acc + addend, or -1 if acc is already -1 or the sum exceeds INT32_MAX. |
Source: `Include/arm_nnsupportfunctions.h:332`
## PACK_S8x4_32x1
`macro` · `c`
```c
#define PACK_S8x4_32x1(v0, v1, v2, v3) ((int32_t)((((uint32_t)(v0)) & 0xFFu) | ((((uint32_t)(v1)) & 0xFFu) << 8) | ((((uint32_t)(v2)) & 0xFFu) << 16) | \ ((((uint32_t)(v3)) & 0xFFu) << 24)))
```
definition to pack four 8 bit values.
Byte lanes are masked and shifted in uint32_t so a negative value never feeds a signed left shift (UB); masking before the shift keeps the same bits the old shift-then-mask form kept. Bit-identical for every input. Deliberate divergence from upstream ARM-software/CMSIS-NN, which still carries the signed-shift form do not paste the upstream text back on a sync (issue #357).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| v0 | | | |
| v1 | | | |
| v2 | | | |
| v3 | | | |
Source: `Include/arm_nnsupportfunctions.h:358`
## PACK_Q15x2_32x1
`macro` · `c`
```c
#define PACK_Q15x2_32x1(v0, v1) ((int32_t)((((uint32_t)(v0)) & 0xFFFFu) | (((uint32_t)(v1)) << 16)))
```
definition to pack two 16 bit values.
Same treatment: the high half is shifted in uint32_t, not int32_t, so a negative v1 is defined; the low half keeps its mask. Bit-identical for every input. Same deliberate upstream divergence as PACK_S8x4_32x1 above.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| v0 | | | |
| v1 | | | |
Source: `Include/arm_nnsupportfunctions.h:369`
## GetNearestNeighbor
`function` · `c`
```c
static int32_t GetNearestNeighbor(
const int input_value,
const int32_t input_size,
const float scale,
const float offset,
const bool align_corners,
const bool half_pixel_centers
)
```
Map an output index to the nearest input index for resize.
This helper follows the TensorFlow Lite nearest-neighbor resize mapping rules.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_value | const int | in | Output index (x or y). |
| input_size | const int32_t | in | Input size along the same axis. |
| scale | const float | in | Precomputed scaling factor for the axis. |
| offset | const float | in | Precomputed offset for the axis. |
| align_corners | const bool | in | If true, use align-corners scaling. |
| half_pixel_centers | const bool | in | If true, use half-pixel center offset. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Nearest input index for the given output index. |
Source: `Include/arm_nnsupportfunctions.h:391`
## arm_nn_is_convolve_1x1
`function` · `c`
```c
static bool arm_nn_is_convolve_1x1(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims
)
```
Check if convolution parameters correspond to a 1x1 convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input dimensions |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | true if parameters describe a 1x1 convolution, false otherwise. |
Source: `Include/arm_nnsupportfunctions.h:416`
## arm_nn_is_convolve_1x1_fast
`function` · `c`
```c
static bool arm_nn_is_convolve_1x1_fast(const cmsis_nn_conv_params *conv_params)
```
Check if a 1x1 convolution qualifies for the fast (unit stride) path.
:::note
Does not validate that the kernel is 1x1. Call arm_nn_is_convolve_1x1() first.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | true if stride is 1x1, false otherwise. |
Source: `Include/arm_nnsupportfunctions.h:432`
## arm_nn_is_convolve_1_x_n
`function` · `c`
```c
static bool arm_nn_is_convolve_1_x_n(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims
)
```
Check if convolution parameters correspond to a 1xN convolution.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input dimensions |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | true if parameters describe a 1xN convolution, false otherwise. |
Source: `Include/arm_nnsupportfunctions.h:444`
## arm_nn_convolve_1_x_n_padding_supported
`function` · `c`
```c
static bool arm_nn_convolve_1_x_n_padding_supported(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Check that `arm_convolve_1_x_n_s4()` handles the horizontal padding of a 1xN convolution.
The kernel places pad.w columns on the left and pad.w + (total_pad % 2) on the right, where total_pad = (output W - 1) * stride.w + filter W - input W, and needs the output columns that read padding to fit in output W. Its padded-column code also assumes that each such column reads at least one input column and that the filter is no wider than the input; otherwise it forms input and filter addresses outside the tensors. On MVE builds another pad placement, or too many padded columns, returns ARM_CMSIS_NN_FAILURE. A VALID layer whose stride leaves trailing input unused (negative total_pad) is therefore rejected, and the wrapper routes it to another convolution. The kernel also computes a single output row, so vertical padding or an output height other than 1 is rejected. A non-positive stride.w is left to the kernel's argument checks.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input dimensions |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | true when `arm_convolve_1_x_n_s4()` handles the padding, false otherwise. |
Source: `Include/arm_nnsupportfunctions.h:473`
## arm_nn_convolve_1_x_n_s8_padding_supported
`function` · `c`
```c
static bool arm_nn_convolve_1_x_n_s8_padding_supported(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Check that `arm_convolve_1_x_n_s8()` accepts the padding and output shape of a 1xN convolution.
The kernel computes a single output row for any pad.w >= 0 and any output width, including an odd total padding, a filter wider than the input and a VALID layer whose stride leaves trailing input unused. It rejects vertical padding, an output height other than 1, a negative pad.w and an empty filter; the wrapper routes those layers to another convolution. A non-positive stride.w is left to the kernel's argument checks.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | true when `arm_convolve_1_x_n_s8()` computes the layer, false otherwise. |
Source: `Include/arm_nnsupportfunctions.h:517`
## arm_nn_convolve_1_x_n_padded_columns
`function` · `c`
```c
static void arm_nn_convolve_1_x_n_padded_columns(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int64_t *left_num,
int64_t *right_num
)
```
Count the output columns of a 1xN convolution whose window reads padding.
Output column j reads input columns j * stride.w - pad.w to j * stride.w - pad.w + filter W - 1. The leading columns whose window starts before the input are left-padded; of the others, the trailing columns whose window ends past the input are right-padded. A window can do both only when it is left-padded.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters. stride.w >= 1 and pad.w >= 0. |
| input_dims | const cmsis_nn_dims * | in | Input dimensions. w >= 0. |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. w >= 1. |
| output_dims | const cmsis_nn_dims * | in | Output dimensions. w >= 0. |
| left_num | int64_t * | out | Number of left-padded output columns, output W at most. |
| right_num | int64_t * | out | Number of right-padded output columns, output W - left_num at most. |
Source: `Include/arm_nnsupportfunctions.h:539`
## arm_nn_dw_conv_opt_dilation_supported
`function` · `c`
```c
static bool arm_nn_dw_conv_opt_dilation_supported(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Check if the dilation, stride and padding of a depthwise layer allow the `arm_depthwise_conv_s8_opt()` or `arm_depthwise_conv_fast_s16()` route.
:::note
Does not check ch_mult, the batch count or the kernel size: `arm_depthwise_conv_wrapper_s8()`, `arm_depthwise_conv_wrapper_s16()` and their buffer-size functions apply their own conditions on those, and all of them take this predicate so that routing and sizing agree.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input dimensions |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output dimensions |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | true for an undilated layer (dilation 1 in both dimensions), or for a 1D layer dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, no vertical padding, dilation.h == 1 and dilation.w >= 1. false otherwise. |
Source: `Include/arm_nnsupportfunctions.h:572`
## arm_q7_to_q15_with_offset
`function` · `c`
```c
void arm_q7_to_q15_with_offset(const int8_t *src, int16_t *dst, int32_t block_size, int16_t offset)
```
Converts the elements from a s8 vector to a s16 vector with an added offset.
Output elements are ordered. The equation used for the conversion process is:
dst[n] = (int16_t) src[n] + offset; 0 <= n < block_size.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| src | const int8_t * | in | pointer to the s8 input vector |
| dst | int16_t * | out | pointer to the s16 output vector |
| block_size | int32_t | in | length of the input vector |
| offset | int16_t | in | s16 offset to be added to each input vector element. |
Source: `Include/arm_nnsupportfunctions.h:649`
## arm_depthwise_conv_s8_opt_get_buffer_size_mve
`function` · `c`
```c
int32_t arm_depthwise_conv_s8_opt_get_buffer_size_mve(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with MVE extension.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_s8_opt_get_buffer_size()`. Note also this is a support function, so not recommended to call directly even on Host.
:::
:::note
This leg sizes its buffer from a fixed channel block rather than from input_dims->c, but it applies the same dimension check as `arm_depthwise_conv_s8_opt_get_buffer_size()` anyway, returning -1 for a negative input_dims->c or filter dimension, or for a byte count that would not fit in an int32_t. Without that check this entry point - and every s4 depthwise sizer, which route here - answered a negative channel count with a plausible positive size (issue #318).
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnsupportfunctions.h:695`
## arm_depthwise_conv_s8_opt_get_buffer_size_dsp
`function` · `c`
```c
int32_t arm_depthwise_conv_s8_opt_get_buffer_size_dsp(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```
Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with DSP extension.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.
:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_s8_opt_get_buffer_size()`. Note also this is a support function, so not recommended to call directly even on Host.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Source: `Include/arm_nnsupportfunctions.h:710`
## arm_nn_depthwise_conv_s8_core
`function` · `c`
```c
int8_t * arm_nn_depthwise_conv_s8_core(
const int8_t *row,
const int16_t *col,
const uint16_t num_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t kernel_size,
const int32_t *const output_bias,
int8_t *out
)
```
Depthwise conv on an im2col buffer where the input channel equals output channel.
Supported framework: TensorFlow Lite micro.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| row | const int8_t * | in | pointer to row |
| col | const int16_t * | in | pointer to im2col buffer, always consists of 2 columns. |
| num_ch | const uint16_t | in | number of channels |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| activation_min | const int32_t | in | minimum value to clamp the output to. Range : int8 |
| activation_max | const int32_t | in | maximum value to clamp the output to. Range : int8 |
| kernel_size | const uint16_t | in | number of elements in one column. |
| output_bias | const int32_t *const | in | per output channel bias. Range : int32 |
| out | int8_t * | out | pointer to output |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the two
1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |
Source: `Include/arm_nnsupportfunctions.h:732`
## arm_nn_mat_mult_s8
`function` · `c`
```c
int8_t * arm_nn_mat_mult_s8(
const int8_t *input_row,
const int8_t *input_col,
const uint16_t output_ch,
const uint16_t col_batches,
const int32_t *output_shift,
const int32_t *output_mult,
const int32_t out_offset,
const int32_t col_offset,
const int32_t row_offset,
const int16_t out_activation_min,
const int16_t out_activation_max,
const uint16_t row_len,
const int32_t *const bias,
int8_t *out
)
```
General Matrix-multiplication function with per-channel requantization.
Supported framework: TensorFlow Lite
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_row | const int8_t * | in | pointer to row operand |
| input_col | const int8_t * | in | pointer to col operand |
| output_ch | const uint16_t | in | number of rows of input_row |
| col_batches | const uint16_t | in | number of column batches. Range: 1 to 4 |
| output_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| output_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| col_offset | const int32_t | in | input tensor(col) offset. |
| row_offset | const int32_t | in | kernel offset(row). Not used. |
| out_activation_min | const int16_t | in | minimum value to clamp the output to. Range : int8 |
| out_activation_max | const int16_t | in | maximum value to clamp the output to. Range : int8 |
| row_len | const uint16_t | in | number of elements in each row |
| bias | const int32_t *const | in | per output channel bias. Range : int32 |
| out | int8_t * | in, out | pointer to output |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the two
1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |
Source: `Include/arm_nnsupportfunctions.h:766`
## arm_nn_mat_mult_kernel_s16
`function` · `c`
```c
int16_t * arm_nn_mat_mult_kernel_s16(
const int8_t *input_a,
const int16_t *input_b,
const int32_t output_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t activation_min,
const int32_t activation_max,
const int32_t num_col_a,
const cmsis_nn_bias_data *const bias_data,
int16_t *out_0,
const int32_t row_address_offset
)
```
Matrix-multiplication function for convolution with per-channel requantization for 16 bits convolution.
This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_a | const int8_t * | in | pointer to operand A |
| input_b | const int16_t * | in | pointer to operand B, always consists of 2 vectors. |
| output_ch | const int32_t | in | number of rows of A |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| activation_min | const int32_t | in | minimum value to clamp the output to. Range : int16 |
| activation_max | const int32_t | in | maximum value to clamp the output to. Range : int16 |
| num_col_a | const int32_t | in | number of columns of A |
| bias_data | const cmsis_nn_bias_data *const | in | pointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct. |
| out_0 | int16_t * | in, out | pointer to output |
| row_address_offset | const int32_t | in | Address offset between rows in output. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the two
1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |
Source: `Include/arm_nnsupportfunctions.h:804`
## arm_nn_mat_mul_core_1x_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mul_core_1x_s8(
int32_t row_elements,
const int32_t skipped_row_elements,
const int8_t *row_base_ref,
const int8_t *col_base_ref,
const int32_t out_ch,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const int32_t *bias,
int8_t *output
)
```
General Vector by Matrix multiplication with requantization and storage of result.
Pseudo-code *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in 'output'
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| row_elements | int32_t | in | number of row elements |
| skipped_row_elements | const int32_t | in | number of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch |
| row_base_ref | const int8_t * | in | pointer to row operand |
| col_base_ref | const int8_t * | in | pointer to col operand |
| out_ch | const int32_t | in | Number of output channels |
| conv_params | const cmsis_nn_conv_params * | in | Pointer to convolution parameters like offsets and activation values |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Pointer to per-channel quantization parameters |
| bias | const int32_t * | in | Pointer to optional per-channel bias |
| output | int8_t * | out | Pointer to output where int8 results are stored. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory. |
Source: `Include/arm_nnsupportfunctions.h:843`
## arm_nn_mat_mul_core_1x_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mul_core_1x_s4(
int32_t row_elements,
const int32_t skipped_row_elements,
const int8_t *row_base_ref,
const int8_t *col_base_ref,
const int32_t out_ch,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const int32_t *bias,
int8_t *output
)
```
General Vector by Matrix multiplication with requantization, storage of result and int4 weights packed into an int8 buffer.
Pseudo-code as int8 example. Int4 filter data will be unpacked. *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in 'output'
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| row_elements | int32_t | in | number of row elements |
| skipped_row_elements | const int32_t | in | number of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch |
| row_base_ref | const int8_t * | in | pointer to row operand |
| col_base_ref | const int8_t * | in | pointer to col operand as packed int4 |
| out_ch | const int32_t | in | Number of output channels |
| conv_params | const cmsis_nn_conv_params * | in | Pointer to convolution parameters like offsets and activation values |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Pointer to per-channel quantization parameters |
| bias | const int32_t * | in | Pointer to optional per-channel bias |
| output | int8_t * | out | Pointer to output where int8 results are stored. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory. |
Source: `Include/arm_nnsupportfunctions.h:881`
## arm_nn_mat_mul_core_4x_s8
`function` · `c`
```c
int8_t * arm_nn_mat_mul_core_4x_s8(
const int32_t row_elements,
const int32_t offset,
const int8_t *row_base,
const int8_t *col_base,
const int32_t out_ch,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const int32_t *bias,
int8_t *output
)
```
Matrix-multiplication with requantization & activation function for four rows and one column.
Compliant to TFLM int8 specification. MVE implementation only
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| row_elements | const int32_t | in | number of row elements |
| offset | const int32_t | in | offset between rows. Can be the same as row_elements. For e.g, in a 1x1 conv scenario with stride as 1. |
| row_base | const int8_t * | in | pointer to row operand |
| col_base | const int8_t * | in | pointer to col operand |
| out_ch | const int32_t | in | Number of output channels |
| conv_params | const cmsis_nn_conv_params * | in | Pointer to convolution parameters like offsets and activation values |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Pointer to per-channel quantization parameters |
| bias | const int32_t * | in | Pointer to per-channel bias |
| output | int8_t * | out | Pointer to output where int8 results are stored. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns the updated output pointer or NULL if implementation is not available. |
Source: `Include/arm_nnsupportfunctions.h:908`
## arm_nn_mat_mult_nt_t_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s4(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *bias,
int8_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t lhs_cols_offset
)
```
General Matrix-multiplication function with per-channel requantization. This function assumes:
- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)
- RHS is int8 packed with 2x int4
- LHS is int8
:::note
This operation also performs the broadcast bias addition before the requantization
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias | const int32_t * | in | Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst | int8_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_offset | const int32_t | in | Offset to be applied the output result |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int8 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int8 |
| lhs_cols_offset | const int32_t | in | Column offset between subsequent lhs_rows |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:950`
## arm_nn_mat_mult_nt_interleaved_t_even_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_interleaved_t_even_s4(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *bias,
int8_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t lhs_cols_offset
)
```
General Matrix-multiplication function with per-channel requantization. This function assumes:
- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)
- RHS is int8 packed with 2x int4
- LHS is int8
- LHS/RHS input columns must be even numbered
- LHS must be interleaved. Compare to arm_nn_mat_mult_nt_t_s4 where LHS is not interleaved.
:::note
This operation also performs the broadcast bias addition before the requantization
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias | const int32_t * | in | Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst | int8_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns. Note this must be even. |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_offset | const int32_t | in | Offset to be applied the output result |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int8 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int8 |
| lhs_cols_offset | const int32_t | in | Column offset between subsequent lhs_rows |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:999`
## arm_nn_mat_mult_nt_t_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s8(
const int32_t *weight_sum_buf,
const int8_t *lhs,
const int8_t *rhs,
const int32_t *bias,
int8_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t row_address_offset,
const int32_t lhs_cols_offset
)
```
General Matrix-multiplication function with per-channel requantization. This function assumes:
- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)
:::note
This operation also performs the broadcast bias addition before the requantization
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| weight_sum_buf | const int32_t * | in | Pointer to the weight sum multiplied by lhs_offset and summed bias buffer |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias | const int32_t * | in | Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst | int8_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_offset | const int32_t | in | Offset to be applied the output result |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int8 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int8 |
| row_address_offset | const int32_t | in | Address offset between rows in output. NOTE: Only used for MVEI extension. |
| lhs_cols_offset | const int32_t | in | Column offset between subsequent lhs_rows |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1046`
## arm_nn_mat_mult_nt_t_1x1_out_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_1x1_out_s8(
const int32_t *weight_sum_buf,
const int8_t *lhs,
const int8_t *rhs,
const int32_t *bias,
int8_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t row_address_offset,
const int32_t lhs_cols_offset
)
```
General Matrix-multiplication function with per-channel requantization. Output is calculated with multiple channels in parallel, rather than multiple output indices in a single channel This function assumes:
- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)
:::note
This operation also performs the broadcast bias addition before the requantization
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| weight_sum_buf | const int32_t * | in | Pointer to the weight sum multiplied by lhs_offset and summed bias buffer |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias | const int32_t * | in | Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst | int8_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_offset | const int32_t | in | Offset to be applied the output result |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int8 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int8 |
| row_address_offset | const int32_t | in | Address offset between rows in output. NOTE: Only used for MVEI extension. |
| lhs_cols_offset | const int32_t | in | Column offset between subsequent lhs_rows |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1097`
## arm_nn_mat_mult_nt_t_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s16(
const int16_t *lhs,
const int8_t *rhs,
const cmsis_nn_bias_data *bias_data,
int16_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t activation_min,
const int32_t activation_max,
const int32_t row_address_offset
)
```
General Matrix-multiplication function with per-channel requantization and int16 input (LHS) and output. This function assumes:
- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)
:::note
This operation also performs the broadcast bias addition before the requantization
:::
MVE implementation only.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias_data | const cmsis_nn_bias_data * | in | Pointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct. |
| dst | int16_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int16 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int16 |
| row_address_offset | const int32_t | in | Address offset between rows in output. NOTE: Only used for MVEI extension. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` or `ARM_CMSIS_NN_NO_IMPL_ERROR` if not for MVE \|---row_address_offset---\| \|____rhs_rows__________________\|
\| \| \| \| --- \| --- \| \| \| \| \| \| \|
\| \| \| lhs_rows
\| \| \| \| --- \| --- \| \| _______________ \| ______________ \| |
Source: `Include/arm_nnsupportfunctions.h:1155`
## arm_nn_mat_mult_nt_t_s8_s32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s8_s32(
const int8_t *lhs,
const int8_t *rhs,
int32_t *dst,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_idx_offset
)
```
General Matrix-multiplication function with int8 input and int32 output. This function assumes:
- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)
:::note
Dst/output buffer must be zeroed out before calling this function.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| dst | int32_t * | in, out | Pointer to the output matrix with "m" rows and "n" columns. Accumulated into, so it must be zeroed by the caller before the call |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of LHS input columns/RHS input rows |
| rhs_cols | const int32_t | in | Number of RHS input columns |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_idx_offset | const int32_t | in | Offset between subsequent output results |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1189`
## arm_nn_vec_mat_mult_t_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s4(
const int8_t *lhs,
const int8_t *packed_rhs,
const int32_t *bias,
int8_t *dst,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)
```
s4 Vector by Matrix (transposed) multiplication
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side vector |
| packed_rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| bias | const int32_t * | in | Input bias |
| dst | int8_t * | out | Output vector |
| lhs_offset | const int32_t | in | Offset to be added to the input values of the left-hand side vector. Range: -127 to 128 |
| dst_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1218`
## arm_nn_vec_mat_mult_t_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s8(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *kernel_sum,
const int32_t *bias,
int8_t *dst,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max,
const int32_t address_offset,
const int32_t rhs_offset
)
```
s8 Vector by Matrix (transposed) multiplication
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| kernel_sum | const int32_t * | in | Kernel sums of the kernels (rhs). See arm_vector_sum_s8 for more info. |
| bias | const int32_t * | in | Input bias |
| dst | int8_t * | out | Output vector |
| lhs_offset | const int32_t | in | Offset to be added to the input values of the left-hand side vector. Range: -127 to 128 |
| dst_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| address_offset | const int32_t | in | Memory position offset for dst. First output is stored at 'dst', the second at 'dst + address_offset' and so on. Default value is typically 1. |
| rhs_offset | const int32_t | in | Offset to be added to the input values of the right-hand side vector. Range: -127 to 128 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1256`
## arm_nn_vec_mat_mult_t_per_ch_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_per_ch_s8(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *kernel_sum,
const int32_t *bias,
int8_t *dst,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t *dst_multiplier,
const int32_t *dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max,
const int32_t address_offset,
const int32_t rhs_offset
)
```
s8 Vector by Matrix (transposed) multiplication using per channel quantization for output
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| kernel_sum | const int32_t * | in | Kernel sums of the kernels (rhs). See arm_vector_sum_s8 for more info. |
| bias | const int32_t * | in | Input bias |
| dst | int8_t * | out | Output vector |
| lhs_offset | const int32_t | in | Offset to be added to the input values of the left-hand side vector. Range: -127 to 128 |
| dst_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| dst_multiplier | const int32_t * | in | Output multipliers |
| dst_shift | const int32_t * | in | Output shifts |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| address_offset | const int32_t | in | Memory position offset for dst. First output is stored at 'dst', the second at 'dst + address_offset' and so on. Default value is typically 1. |
| rhs_offset | const int32_t | in | Offset to be added to the input values of the right-hand side vector. Range: -127 to 128 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1297`
## arm_nn_vec_mat_mult_t_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s16(
const int16_t *lhs,
const int8_t *rhs,
const int64_t *bias,
int16_t *dst,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)
```
s16 Vector by s8 Matrix (transposed) multiplication
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| bias | const int64_t * | in | Input bias |
| dst | int16_t * | out | Output vector |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int16 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1330`
## arm_nn_vec_mat_mult_t_per_ch_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_per_ch_s16(
const int16_t *lhs,
const int8_t *rhs,
const int64_t *bias,
int16_t *dst,
const int32_t *dst_multiplier,
const int32_t *dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)
```
s16 vector(lhs) by s8 matrix (transposed) multiplication and per channel quant output
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| bias | const int64_t * | in | Input bias |
| dst | int16_t * | out | Output vector |
| dst_multiplier | const int32_t * | in | Per channel output multiplier. Length of vector is equal to rhs_rows |
| dst_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to rhs_rows |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int16 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1358`
## arm_nn_vec_mat_mult_t_s16_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s16_s16(
const int16_t *lhs,
const int16_t *rhs,
const int64_t *bias,
int16_t *dst,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)
```
s16 Vector by s16 Matrix (transposed) multiplication
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Input left-hand side vector |
| rhs | const int16_t * | in | Input right-hand side matrix (transposed) |
| bias | const int64_t * | in | Input bias |
| dst | int16_t * | out | Output vector |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int16 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1386`
## arm_nn_vec_mat_mult_t_svdf_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_svdf_s8(
const int8_t *lhs,
const int8_t *rhs,
int16_t *dst,
const int32_t lhs_offset,
const int32_t scatter_offset,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)
```
s8 Vector by Matrix (transposed) multiplication with s16 output
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| dst | int16_t * | out | Output vector |
| lhs_offset | const int32_t | in | Offset to be added to the input values of the left-hand side vector. Range: -127 to 128 |
| scatter_offset | const int32_t | in | Address offset for dst. First output is stored at 'dst', the second at 'dst + scatter_offset' and so on. |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int16 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int16 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:1417`
## arm_nn_depthwise_conv_nt_t_padded_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_padded_s8(
const int8_t *lhs,
const int8_t *rhs,
const int32_t lhs_offset,
const int32_t active_ch,
const int32_t total_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t row_x_col,
const int32_t *const output_bias,
int8_t *out
)
```
Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in padded cases where the padding is -lhs_offset(Range: int8). Dimensions are the same for lhs and rhs.
:::note
Tail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `active_ch`.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side matrix |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| lhs_offset | const int32_t | in | LHS matrix offset(input offset). Range: -127 to 128 |
| active_ch | const int32_t | in | Subset of total_ch processed |
| total_ch | const int32_t | in | Number of channels in LHS/RHS |
| out_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to number of channels |
| out_mult | const int32_t * | in | Per channel output multiplier. Length of vector is equal to number of channels |
| out_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| row_x_col | const uint16_t | in | (row_dimension * col_dimension) of LHS/RHS matrix |
| output_bias | const int32_t *const | in | Per channel output bias. Length of vector is equal to number of channels |
| out | int8_t * | out | Output pointer |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` if an implementation is available or `ARM_CMSIS_NN_NO_IMPL_ERROR` otherwise |
Source: `Include/arm_nnsupportfunctions.h:1453`
## arm_nn_depthwise_conv_nt_t_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_s8(
const int32_t *weight_sum_buf,
const int8_t *lhs,
const int8_t *rhs,
const int32_t lhs_offset,
const int32_t active_ch,
const int32_t total_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t row_x_col,
const int32_t *const output_bias,
int8_t *out
)
```
Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. Dimensions are the same for lhs and rhs.
:::note
Tail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `active_ch`.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| weight_sum_buf | const int32_t * | in | Pointer to the weight sum multiplied by lhs_offset and summed bias buffer |
| lhs | const int8_t * | in | Input left-hand side matrix |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| lhs_offset | const int32_t | in | LHS matrix offset(input offset). Range: -127 to 128 |
| active_ch | const int32_t | in | Subset of total_ch processed |
| total_ch | const int32_t | in | Number of channels in LHS/RHS |
| out_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to number of channels. |
| out_mult | const int32_t * | in | Per channel output multiplier. Length of vector is equal to number of channels. |
| out_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| row_x_col | const uint16_t | in | (row_dimension * col_dimension) of LHS/RHS matrix |
| output_bias | const int32_t *const | in | Per channel output bias. Length of vector is equal to number of channels. |
| out | int8_t * | out | Output pointer |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` if an implementation is available or `ARM_CMSIS_NN_NO_IMPL_ERROR` otherwise |
Source: `Include/arm_nnsupportfunctions.h:1492`
## arm_nn_depthwise_conv_s8_planar_candidate
`function` · `c`
```c
static int32_t arm_nn_depthwise_conv_s8_planar_candidate(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims
)
```
Necessary conditions of the planar rule that are cheap to test inline: at most 32 channels and stride 1. A caller can skip `arm_nn_depthwise_conv_s8_planar()` for layers that fail them without changing which layers it takes.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [1, H, W, C_IN] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 1 when the layer may take the planar path, 0 when it cannot. |
Source: `Include/arm_nnsupportfunctions.h:1517`
## arm_nn_is_convolve_s8_small_cin
`function` · `c`
```c
static int32_t arm_nn_is_convolve_s8_small_cin(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
const cmsis_nn_dims *upscale_dims
)
```
The gate of `arm_convolve_s8_small_cin()`: upscale_dims NULL, input depth 1 to 3 with filter depth equal to it, dilation 1, a kernel of at least 1x1 with kernel width x depth at most 16 and at most 48 values, and a positive multiple of 4 output channels. Plain C; it evaluates the same on every build.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions, or NULL |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 1 when the layer is in the gate, 0 otherwise. |
Source: `Include/arm_nnsupportfunctions.h:1536`
## arm_nn_is_convolve_s8_3x3_c16_s1
`function` · `c`
```c
static int32_t arm_nn_is_convolve_s8_3x3_c16_s1(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *upscale_dims
)
```
The gate of `arm_convolve_s8_3x3_c16_s1()`: upscale_dims NULL, input and filter depth 16, a 3x3 kernel, and stride and dilation 1. Plain C; it evaluates the same on every build.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
| upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions, or NULL |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 1 when the layer is in the gate, 0 otherwise. |
Source: `Include/arm_nnsupportfunctions.h:1562`
## arm_nn_convolve_s8_groups_invalid
`function` · `c`
```c
static int32_t arm_nn_convolve_s8_groups_invalid(
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
The group check of `arm_convolve_s8()`, for its direct entries: with groups = C_IN / filter C, C_IN or C_OUT is not a multiple of groups. A filter C of zero or above C_IN gives no group count and is not reported.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 1 when `arm_convolve_s8()` reports the group count as an argument error, 0 otherwise. |
Source: `Include/arm_nnsupportfunctions.h:1582`
## arm_nn_depthwise_conv_s8_planar_bytes
`function` · `c`
```c
int32_t arm_nn_depthwise_conv_s8_planar_bytes(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)
```
Plane size in bytes that `arm_nn_depthwise_conv_s8_planar()` needs for a layer, or -1 when the layer is not one it takes. The rule is plain C and evaluates the same on every build.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [1, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The plane size in bytes, or -1. |
Source: `Include/arm_nnsupportfunctions.h:1601`
## arm_nn_depthwise_conv_s8_planar
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_depthwise_conv_s8_planar(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input,
const cmsis_nn_dims *filter_dims,
const int8_t *kernel,
const cmsis_nn_dims *output_dims,
int8_t *output
)
```
s8 depthwise convolution with channel multiplier 1 and stride 1, vectorized across the output pixels of one channel plane instead of across channels. It serves the few-channel and 1xk layers of `arm_depthwise_conv_s8_opt()`, with the same scratch buffer and weight sums.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Scratch buffer of `arm_depthwise_conv_s8_opt_get_buffer_size()` bytes |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums from `arm_depthwise_convolve_weight_sum()`, bias included |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [1, H, W, C_IN] |
| input | const int8_t * | in | Input data pointer |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| kernel | const int8_t * | in | Filter data pointer |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
| output | int8_t * | out | Output data pointer |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | `ARM_CMSIS_NN_SUCCESS` when the layer was computed, or `ARM_CMSIS_NN_NO_IMPL_ERROR` when it is not one this path takes or its plane does not fit in ctx->size (then nothing is written), or MVE is not available. |
Source: `Include/arm_nnsupportfunctions.h:1626`
## arm_nn_depthwise_conv_nt_t_s4
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_s4(
const int8_t *lhs,
const int8_t *rhs,
const int32_t lhs_offset,
const int32_t active_ch,
const int32_t total_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t row_x_col,
const int32_t *const output_bias,
int8_t *out
)
```
Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. rhs consists of packed int4 data. Dimensions are the same for lhs and rhs.
:::note
Tail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `active_ch`.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side matrix |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed). Consists of int4 data packed in an int8 buffer. |
| lhs_offset | const int32_t | in | LHS matrix offset(input offset). Range: -127 to 128 |
| active_ch | const int32_t | in | Subset of total_ch processed |
| total_ch | const int32_t | in | Number of channels in LHS/RHS |
| out_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to number of channels. |
| out_mult | const int32_t * | in | Per channel output multiplier. Length of vector is equal to number of channels. |
| out_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| row_x_col | const uint16_t | in | (row_dimension * col_dimension) of LHS/RHS matrix |
| output_bias | const int32_t *const | in | Per channel output bias. Length of vector is equal to number of channels. |
| out | int8_t * | out | Output pointer |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the two
- Updated output pointer if an implementation is available - NULL if no implementation is available. |
Source: `Include/arm_nnsupportfunctions.h:1663`
## arm_nn_depthwise_conv_nt_t_s16
`function` · `c`
```c
int16_t * arm_nn_depthwise_conv_nt_t_s16(
const int16_t *lhs,
const int8_t *rhs,
const uint16_t num_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t row_x_col,
const int64_t *const output_bias,
int16_t *out
)
```
Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. Dimensions are the same for lhs and rhs.
:::note
Tail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `num_ch`.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Input left-hand side matrix |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| num_ch | const uint16_t | in | Number of channels in LHS/RHS |
| out_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to number of channels. |
| out_mult | const int32_t * | in | Per channel output multiplier. Length of vector is equal to number of channels. |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| row_x_col | const uint16_t | in | (row_dimension * col_dimension) of LHS/RHS matrix |
| output_bias | const int64_t *const | in | Per channel output bias. Length of vector is equal to number of channels. |
| out | int16_t * | out | Output pointer |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the two
- Updated output pointer if an implementation is available - NULL if no implementation is available. |
Source: `Include/arm_nnsupportfunctions.h:1699`
## arm_nn_transpose_conv_row_s8_s32
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_transpose_conv_row_s8_s32(
const int8_t *lhs,
const int8_t *rhs,
int32_t *output_start,
const int32_t output_index,
const int32_t output_max,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t input_channels,
const int32_t output_channels,
const int32_t lhs_offset,
const int32_t row_offset,
const int32_t input_x,
const int32_t stride_x,
const int32_t skip_row_top,
const int32_t skip_row_bottom
)
```
Row of s8 scalars multiplicated with a s8 matrix ad accumulated into a s32 rolling scratch buffer. Helpfunction for transposed convolution.
:::note
Rolling buffer refers to how the function wraps around the scratch buffer, e.g. it starts writing at [output_start + output_index], writes to [output_start + output_max] and then continues at [output_start] again.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side scalars |
| rhs | const int8_t * | in | Input right-hand side matrix |
| output_start | int32_t * | out | Output buffer start |
| output_index | const int32_t | in | Output buffer current index |
| output_max | const int32_t | in | Output buffer size |
| rhs_rows | const int32_t | in | Number of rows in rhs matrix |
| rhs_cols | const int32_t | in | Number of columns in rhs matrix |
| input_channels | const int32_t | in | Number of input channels |
| output_channels | const int32_t | in | Number of output channels |
| lhs_offset | const int32_t | in | Offset added to lhs before multiplication |
| row_offset | const int32_t | in | Address offset between each row of data output |
| input_x | const int32_t | in | Length of lhs scalar row. |
| stride_x | const int32_t | in | Address offset between each scalar-matrix multiplication result. |
| skip_row_top | const int32_t | in | Skip rows on top of the filter, used for padding. |
| skip_row_bottom | const int32_t | in | Skip rows in the bottom of the filter, used for padding. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnsupportfunctions.h:1735`
## arm_nn_read_q15x2_ia
`function` · `c`
```c
static int32_t arm_nn_read_q15x2_ia(const int16_t **in_q15)
```
Read 2 s16 elements and post increment pointer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_q15 | const int16_t ** | in, out | Pointer to pointer that holds address of input. Advanced past the elements read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | q31 value |
Source: `Include/arm_nnsupportfunctions.h:1756`
## arm_nn_read_s8x4_ia
`function` · `c`
```c
static int32_t arm_nn_read_s8x4_ia(const int8_t **in_s8)
```
Read 4 s8 from s8 pointer and post increment pointer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_s8 | const int8_t ** | in, out | Pointer to pointer that holds address of input. Advanced past the elements read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | q31 value |
Source: `Include/arm_nnsupportfunctions.h:1771`
## arm_nn_read_s8x2_ia
`function` · `c`
```c
static int32_t arm_nn_read_s8x2_ia(const int8_t **in_s8)
```
Read 2 s8 from s8 pointer and post increment pointer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_s8 | const int8_t ** | in, out | Pointer to pointer that holds address of input. Advanced past the elements read. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | q31 value |
Source: `Include/arm_nnsupportfunctions.h:1785`
## arm_nn_read_s16x2
`function` · `c`
```c
static int32_t arm_nn_read_s16x2(const int16_t *in)
```
Read 2 int16 values from int16 pointer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in | const int16_t * | in | pointer to address of input. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | s32 value |
Source: `Include/arm_nnsupportfunctions.h:1799`
## arm_nn_read_s8x4
`function` · `c`
```c
static int32_t arm_nn_read_s8x4(const int8_t *in_s8)
```
Read 4 s8 values.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_s8 | const int8_t * | in | pointer to address of input. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | s32 value |
Source: `Include/arm_nnsupportfunctions.h:1812`
## arm_nn_read_s8x2
`function` · `c`
```c
static int32_t arm_nn_read_s8x2(const int8_t *in_s8)
```
Read 2 s8 values.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_s8 | const int8_t * | in | pointer to address of input. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | s32 value |
Source: `Include/arm_nnsupportfunctions.h:1824`
## arm_nn_write_s8x4_ia
`function` · `c`
```c
static void arm_nn_write_s8x4_ia(int8_t **in, int32_t value)
```
Write four s8 to s8 pointer and increment pointer afterwards.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in | int8_t ** | in, out | Double pointer to destination. Advanced past the bytes written. |
| value | int32_t | in | Four bytes to copy |
Source: `Include/arm_nnsupportfunctions.h:1837`
## arm_memset_s8
`function` · `c`
```c
static void arm_memset_s8(int8_t *dst, const int8_t val, uint32_t block_size)
```
memset optimized for MVE
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int8_t * | in, out | Destination pointer |
| val | const int8_t | in | Value to set |
| block_size | uint32_t | in | Number of bytes to copy. |
Source: `Include/arm_nnsupportfunctions.h:1850`
## arm_memset_s16
`function` · `c`
```c
static void arm_memset_s16(int16_t *dst, const int16_t val, uint32_t block_size)
```
memset optimized for MVE for 16-bit data.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int16_t * | in, out | Destination pointer. |
| val | const int16_t | in | 16-bit value to set. |
| block_size | uint32_t | in | Number of int16_t values to set. |
Source: `Include/arm_nnsupportfunctions.h:1873`
## arm_nn_mat_mult_kernel_s4_s16
`function` · `c`
```c
int8_t * arm_nn_mat_mult_kernel_s4_s16(
const int8_t *input_a,
const int16_t *input_b,
const uint16_t output_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t num_col_a,
const int32_t *const output_bias,
int8_t *out_0
)
```
Matrix-multiplication function for convolution with per-channel requantization and 4 bit weights.
This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_a | const int8_t * | in | pointer to operand A, int8 packed with 2x int4. |
| input_b | const int16_t * | in | pointer to operand B, always consists of 2 vectors. |
| output_ch | const uint16_t | in | number of rows of A |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| activation_min | const int32_t | in | minimum value to clamp the output to. Range : int8 |
| activation_max | const int32_t | in | maximum value to clamp the output to. Range : int8 |
| num_col_a | const int32_t | in | number of columns of A |
| output_bias | const int32_t *const | in | per output channel bias. Range : int32 |
| out_0 | int8_t * | in, out | pointer to output |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the two
1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |
Source: `Include/arm_nnsupportfunctions.h:2094`
## arm_nn_mat_mult_kernel_s8_s16
`function` · `c`
```c
int8_t * arm_nn_mat_mult_kernel_s8_s16(
const int8_t *input_a,
const int16_t *input_b,
const uint16_t output_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int16_t activation_min,
const int16_t activation_max,
const int32_t num_col_a,
const int32_t aligned_num_col_a,
const int32_t *const output_bias,
int8_t *out_0
)
```
Matrix-multiplication function for convolution with per-channel requantization.
This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_a | const int8_t * | in | pointer to operand A |
| input_b | const int16_t * | in | pointer to operand B, always consists of 2 vectors. |
| output_ch | const uint16_t | in | number of rows of A |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| activation_min | const int16_t | in | minimum value to clamp the output to. Range : int8 |
| activation_max | const int16_t | in | maximum value to clamp the output to. Range : int8 |
| num_col_a | const int32_t | in | number of columns of A |
| aligned_num_col_a | const int32_t | in | number of columns of A aligned by 4 |
| output_bias | const int32_t *const | in | per output channel bias. Range : int32 |
| out_0 | int8_t * | in, out | pointer to output |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the two
1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |
Source: `Include/arm_nnsupportfunctions.h:2128`
## arm_nn_mat_mult_kernel_row_offset_s8_s16
`function` · `c`
```c
int8_t * arm_nn_mat_mult_kernel_row_offset_s8_s16(
const int8_t *input_a,
const int16_t *input_b,
const uint16_t output_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int16_t activation_min,
const int16_t activation_max,
const int32_t num_col_a,
const int32_t aligned_num_col_a,
const int32_t *const output_bias,
const int32_t row_address_offset,
int8_t *out_0
)
```
Matrix-multiplication function for convolution with per-channel requantization, supporting an address offset between rows.
This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max.
This function is slighly less performant than arm_nn_mat_mult_kernel_s8_s16, but allows support for grouped convolution. Supported framework: TensorFlow Lite micro.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_a | const int8_t * | in | pointer to operand A |
| input_b | const int16_t * | in | pointer to operand B, always consists of 2 vectors. |
| output_ch | const uint16_t | in | number of rows of A |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| activation_min | const int16_t | in | minimum value to clamp the output to. Range : int8 |
| activation_max | const int16_t | in | maximum value to clamp the output to. Range : int8 |
| num_col_a | const int32_t | in | number of columns of A |
| aligned_num_col_a | const int32_t | in | number of columns of A aligned by 4 |
| output_bias | const int32_t *const | in | per output channel bias. Range : int32 |
| row_address_offset | const int32_t | in | address offset between rows in the output |
| out_0 | int8_t * | in, out | pointer to output |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns one of the two
1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |
Source: `Include/arm_nnsupportfunctions.h:2168`
## arm_nn_softmax_common_s8
`function` · `c`
```c
void arm_nn_softmax_common_s8(
const int8_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const int32_t diff_min,
const bool int16_output,
void *output
)
```
Common softmax function for s8 input and s8 or s16 output.
:::note
Supported framework: TensorFlow Lite micro (bit-accurate)
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input tensor |
| num_rows | const int32_t | in | Number of rows in the input tensor |
| row_size | const int32_t | in | Number of elements in each input row |
| mult | const int32_t | in | Input quantization multiplier |
| shift | const int32_t | in | Input quantization shift within the range [0, 31] |
| diff_min | const int32_t | in | Minimum difference with max in row. Used to check if the quantized exponential operation can be performed |
| int16_output | const bool | in | Indicating s8 output if 0 else s16 output |
| output | void * | out | Pointer to the output tensor |
Source: `Include/arm_nnsupportfunctions.h:2197`
## NN_ROUND
`macro` · `c`
```c
#define NN_ROUND(out_shift) ((0x1 << out_shift) >> 1)
```
macro for adding rounding offset
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| out_shift | | | |
Source: `Include/arm_nnsupportfunctions.h:2210`
## MUL_SAT
`macro` · `c`
```c
#define MUL_SAT(a, b) arm_nn_doubling_high_mult((a), (b))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | | | |
| b | | | |
Source: `Include/arm_nnsupportfunctions.h:2216`
## MUL_SAT_MVE
`macro` · `c`
```c
#define MUL_SAT_MVE(a, b) arm_doubling_high_mult_mve_32x4((a), (b))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | | | |
| b | | | |
Source: `Include/arm_nnsupportfunctions.h:2217`
## MUL_POW2
`macro` · `c`
```c
#define MUL_POW2(a, b) arm_nn_mult_by_power_of_two((a), (b))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | | | |
| b | | | |
Source: `Include/arm_nnsupportfunctions.h:2218`
## DIV_POW2
`macro` · `c`
```c
#define DIV_POW2(a, b) arm_nn_divide_by_power_of_two((a), (b))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | | | |
| b | | | |
Source: `Include/arm_nnsupportfunctions.h:2220`
## DIV_POW2_MVE
`macro` · `c`
```c
#define DIV_POW2_MVE(a, b) arm_divide_by_power_of_two_mve((a), (b))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | | | |
| b | | | |
Source: `Include/arm_nnsupportfunctions.h:2221`
## EXP_ON_NEG
`macro` · `c`
```c
#define EXP_ON_NEG(x) arm_nn_exp_on_negative_values((x))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | | | |
Source: `Include/arm_nnsupportfunctions.h:2223`
## ONE_OVER1
`macro` · `c`
```c
#define ONE_OVER1(x) arm_nn_one_over_one_plus_x_for_x_in_0_1((x))
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | | | |
Source: `Include/arm_nnsupportfunctions.h:2224`
## arm_nn_doubling_high_mult
`function` · `c`
```c
static int32_t arm_nn_doubling_high_mult(const int32_t m1, const int32_t m2)
```
Saturating doubling high multiply. Result matches NEON instruction VQRDMULH.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| m1 | const int32_t | in | Multiplicand. Range: {NN_Q31_MIN, NN_Q31_MAX} |
| m2 | const int32_t | in | Multiplier. Range: {NN_Q31_MIN, NN_Q31_MAX} |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Result of multiplication. |
Source: `Include/arm_nnsupportfunctions.h:2234`
## arm_nn_doubling_high_mult_no_sat
`function` · `c`
```c
static int32_t arm_nn_doubling_high_mult_no_sat(int32_t m1, int32_t m2)
```
Doubling high multiply without saturation. This is intended for requantization where the scale is a positive integer.
:::note
The result of this matches that of neon instruction VQRDMULH for m1 in range {NN_Q31_MIN, NN_Q31_MAX} and m2 in range {NN_Q31_MIN + 1, NN_Q31_MAX}. Saturation occurs when m1 equals m2 equals NN_Q31_MIN and that is not handled by this function.
:::
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| m1 | int32_t | in | Multiplicand. Range: {NN_Q31_MIN, NN_Q31_MAX} |
| m2 | int32_t | in | Multiplier Range: {NN_Q31_MIN, NN_Q31_MAX} |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Result of multiplication. |
Source: `Include/arm_nnsupportfunctions.h:2272`
## arm_nn_divide_by_power_of_two
`function` · `c`
```c
static int32_t arm_nn_divide_by_power_of_two(const int32_t dividend, const int32_t exponent)
```
Rounding divide by power of two.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dividend | const int32_t | in | - Dividend |
| exponent | const int32_t | in | - Divisor = power(2, exponent) Range: [0, 31] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Rounded result of division. Midpoint is rounded away from zero. |
Source: `Include/arm_nnsupportfunctions.h:2323`
## arm_nn_nonneg_divide_by_pot_s32
`function` · `c`
```c
static int32_t arm_nn_nonneg_divide_by_pot_s32(int32_t dividend, int32_t exponent)
```
Rounding divide by power of two for non-negative values.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dividend | int32_t | in | - Dividend (assumed to be non-negative) |
| exponent | int32_t | in | - Divisor = power(2, exponent) Range: [0, 31] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Rounded result of division. Midpoint is rounded away from zero. |
Source: `Include/arm_nnsupportfunctions.h:2378`
## arm_nn_requantize
`function` · `c`
```c
static int32_t arm_nn_requantize(const int32_t val, const int32_t multiplier, const int32_t shift)
```
Requantize a given value.
Essentially returns (val * multiplier)/(2 ^ shift) with different rounding depending if CMSIS_NN_USE_SINGLE_ROUNDING is defined or not.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | const int32_t | in | Value to be requantized |
| multiplier | const int32_t | in | Multiplier. Range {NN_Q31_MIN + 1, Q32_MAX} |
| shift | const int32_t | in | Shift. Range: {-31, 30} Default branch: If shift is positive left shift 'val * multiplier' with shift If shift is negative right shift 'val * multiplier' with abs(shift) Single round branch: Input for total_shift in divide by '2 ^ total_shift' |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Default branch: Returns (val * multiplier) with rounding divided by (2 ^ shift) with rounding Single round branch: Returns (val * multiplier)/(2 ^ (31 - shift)) with rounding |
Source: `Include/arm_nnsupportfunctions.h:2416`
## arm_nn_requantize_s64
`function` · `c`
```c
static int32_t arm_nn_requantize_s64(const int64_t val, const int32_t reduced_multiplier, const int32_t shift)
```
Requantize a given 64 bit value.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | const int64_t | in | Value to be requantized in the range {-(1<<47)} to {(1<<47) - 1} |
| reduced_multiplier | const int32_t | in | Reduced multiplier in the range {NN_Q31_MIN + 1, Q32_MAX} to {Q16_MIN + 1, Q16_MAX} |
| shift | const int32_t | in | Left or right shift for 'val * multiplier' in the range {-31} to {7} |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Returns (val * multiplier)/(2 ^ shift) |
Source: `Include/arm_nnsupportfunctions.h:2453`
## arm_nn_sat_lshift_s16
`function` · `c`
```c
static int16_t arm_nn_sat_lshift_s16(int16_t x, int shift)
```
Saturating left shift for int16_t.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | int16_t | in | value to be shifted |
| shift | int | in | Nonpositive values return x; positive values multiply by 2^shift with s16 saturation. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | shifted value |
Source: `Include/arm_nnsupportfunctions.h:2471`
## arm_nn_sqrdmulh_s16
`function` · `c`
```c
static int16_t arm_nn_sqrdmulh_s16(int16_t a, int16_t b)
```
Saturating *Rounding* Doubling High Mul (s16).
Matches NEON SQRDMULH s16
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | int16_t | in | Multiplicand |
| b | int16_t | in | Multiplier |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Result of multiplication. |
Source: `Include/arm_nnsupportfunctions.h:2490`
## arm_nn_sqdmulh_s16
`function` · `c`
```c
static int16_t arm_nn_sqdmulh_s16(int16_t a, int16_t b)
```
Saturating **Non-rounded** Doubling High Mul (s16).
Matches NEON SQDMULH s16
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | int16_t | in | Multiplicand |
| b | int16_t | in | Multiplier |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Result of multiplication. |
Source: `Include/arm_nnsupportfunctions.h:2510`
## arm_nn_divide_by_power_of_two_s16
`function` · `c`
```c
static int16_t arm_nn_divide_by_power_of_two_s16(int16_t x, int exponent)
```
Rounding divide by power of two (s16), midpoint away from zero.
Mirrors arm_nn_divide_by_power_of_two() semantics for s16.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | int16_t | in | Dividend |
| exponent | int | in | Divisor = power(2, exponent) Range: [0, 15] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Rounded result of division. Midpoint is rounded away from zero. |
Source: `Include/arm_nnsupportfunctions.h:2530`
## arm_memcpy_s8
`function` · `c`
```c
static void arm_memcpy_s8(int8_t *dst, const int8_t *src, uint32_t block_size)
```
memcpy optimized for MVE
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int8_t * | in, out | Destination pointer |
| src | const int8_t * | in | Source pointer. |
| block_size | uint32_t | in | Number of bytes to copy. |
Source: `Include/arm_nnsupportfunctions.h:2545`
## arm_memcpy_s16
`function` · `c`
```c
static void arm_memcpy_s16(int16_t *dst, const int16_t *src, uint32_t block_size)
```
memcpy optimized for MVE
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int16_t * | in, out | Destination pointer |
| src | const int16_t * | in | Source pointer. |
| block_size | uint32_t | in | Number of values to copy. |
Source: `Include/arm_nnsupportfunctions.h:2569`
## arm_memcpy_s32
`function` · `c`
```c
static void arm_memcpy_s32(int32_t *dst, const int32_t *src, uint32_t block_size)
```
memcpy optimized for MVE
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int32_t * | in, out | Destination pointer |
| src | const int32_t * | in | Source pointer. |
| block_size | uint32_t | in | Number of values to copy. |
Source: `Include/arm_nnsupportfunctions.h:2581`
## arm_memcpy_q15
`function` · `c`
```c
static void arm_memcpy_q15(int16_t *dst, const int16_t *src, uint32_t block_size)
```
memcpy wrapper for int16
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int16_t * | in, out | Destination pointer |
| src | const int16_t * | in | Source pointer. |
| block_size | uint32_t | in | Number of bytes to copy. |
Source: `Include/arm_nnsupportfunctions.h:2593`
## arm_nn_exp_on_negative_values
`function` · `c`
```c
static int32_t arm_nn_exp_on_negative_values(int32_t val)
```
Fixed-point exp() of a non-positive value.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | int32_t | in | Input in Q5.26 fixed point. Must be less than or equal to 0 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | exp(val) in Q0.31 fixed point. Returns NN_Q31_MAX when `val` is 0. |
Source: `Include/arm_nnsupportfunctions.h:2865`
## SELECT_IF_NON_ZERO
`macro` · `c`
```c
#define SELECT_IF_NON_ZERO(x) { \ mask = MASK_IF_NON_ZERO(remainder & (1 << shift++)); \ result = SELECT_USING_MASK(mask, MUL_SAT(result, x), result); \ }
```
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | | | |
Source: `Include/arm_nnsupportfunctions.h:2878`
## arm_nn_mult_by_power_of_two
`function` · `c`
```c
static int32_t arm_nn_mult_by_power_of_two(const int32_t val, const int32_t exp)
```
Saturating multiply by a power of two.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | const int32_t | in | Value to be multiplied |
| exp | const int32_t | in | Exponent. Multiplier = power(2, exp) |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | val * 2^exp saturated to the int32 range |
Source: `Include/arm_nnsupportfunctions.h:2905`
## arm_nn_one_over_one_plus_x_for_x_in_0_1
`function` · `c`
```c
static int32_t arm_nn_one_over_one_plus_x_for_x_in_0_1(int32_t val)
```
Fixed-point 1 / (1 + x) for x in [0, 1), computed with Newton-Raphson iterations.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | int32_t | in | x in Q0.31 fixed point. Range: [0, NN_Q31_MAX] |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 1 / (1 + x) in Q0.31 fixed point |
Source: `Include/arm_nnsupportfunctions.h:2920`
## arm_nn_write_q15x2_ia
`function` · `c`
```c
static void arm_nn_write_q15x2_ia(int16_t **dest_q15, int32_t src_q31)
```
Write 2 s16 elements and post increment pointer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dest_q15 | int16_t ** | in, out | Pointer to pointer that holds address of destination. Advanced past the elements written. |
| src_q31 | int32_t | in | Input value to be written. |
Source: `Include/arm_nnsupportfunctions.h:2942`
## arm_nn_write_s8x2_ia
`function` · `c`
```c
static void arm_nn_write_s8x2_ia(int8_t **dst, int16_t src)
```
Write 2 s8 elements and post increment pointer.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int8_t ** | in, out | Pointer to pointer that holds address of destination. Advanced past the elements written. |
| src | int16_t | in | Input value to be written. |
Source: `Include/arm_nnsupportfunctions.h:2955`
## arm_cmsis_nn_dim_at
`function` · `c`
```c
static int32_t arm_cmsis_nn_dim_at(const cmsis_nn_dims *dims, int32_t index)
```
Get dimension value at specific index.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dims | const cmsis_nn_dims * | in | Pointer to `cmsis_nn_dims` structure |
| index | int32_t | in | Index of dimension to get |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Dimension value at specified index |
Source: `Include/arm_nnsupportfunctions.h:2969`
## arm_cmsis_nn_shape_product
`function` · `c`
```c
static size_t arm_cmsis_nn_shape_product(const int32_t *shape, int32_t length)
```
Calculate the product of all dimensions in a shape array.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| shape | const int32_t * | in | Pointer to array containing shape dimensions |
| length | int32_t | in | Number of dimensions in the shape array |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | Product of all dimensions |
Source: `Include/arm_nnsupportfunctions.h:2994`
## arm_nn_lstm_step_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_lstm_step_s8(
const int8_t *data_in,
const int8_t *hidden_in,
int8_t *hidden_out,
const cmsis_nn_lstm_params *params,
cmsis_nn_lstm_context *buffers,
const int32_t batch_offset
)
```
Update LSTM function for an iteration step using s8 input and output, and s16 internally.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const int8_t * | in | Data input pointer |
| hidden_in | const int8_t * | in | Hidden state/ recurrent input pointer |
| hidden_out | int8_t * | out | Hidden state/ recurrent output pointer |
| params | const cmsis_nn_lstm_params * | in | Struct containg all information about the lstm operator, see arm_nn_types. |
| buffers | cmsis_nn_lstm_context * | in, out | Struct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnsupportfunctions.h:3022`
## arm_nn_lstm_step_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_lstm_step_s16(
const int16_t *data_in,
const int16_t *hidden_in,
int16_t *hidden_out,
const cmsis_nn_lstm_params *params,
cmsis_nn_lstm_context *buffers,
const int32_t batch_offset
)
```
Update LSTM function for an iteration step using s16 input and output, and s16 internally.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const int16_t * | in | Data input pointer |
| hidden_in | const int16_t * | in | Hidden state/ recurrent input pointer |
| hidden_out | int16_t * | out | Hidden state/ recurrent output pointer |
| params | const cmsis_nn_lstm_params * | in | Struct containg all information about the lstm operator, see arm_nn_types. |
| buffers | cmsis_nn_lstm_context * | in, out | Struct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnsupportfunctions.h:3046`
## arm_nn_lstm_calculate_gate_s8_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_lstm_calculate_gate_s8_s16(
const int8_t *data_in,
const int8_t *hidden_in,
const cmsis_nn_lstm_gate *gate_data,
const cmsis_nn_lstm_params *params,
int16_t *output,
const int32_t batch_offset
)
```
Updates a LSTM gate for an iteration step of LSTM function, int8x8_16 version.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const int8_t * | in | Data input pointer |
| hidden_in | const int8_t * | in | Hidden state/ recurrent input pointer |
| gate_data | const cmsis_nn_lstm_gate * | in | Struct containing all information about the gate caluclation, see arm_nn_types. |
| params | const cmsis_nn_lstm_params * | in | Struct containing all information about the lstm_operation, see arm_nn_types |
| output | int16_t * | out | Hidden state/ recurrent output pointer |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches, see arm_nn_lstm_step_s8. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnsupportfunctions.h:3067`
## arm_nn_lstm_calculate_gate_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_lstm_calculate_gate_s16(
const int16_t *data_in,
const int16_t *hidden_in,
const cmsis_nn_lstm_gate *gate_data,
const cmsis_nn_lstm_params *params,
int16_t *output,
const int32_t batch_offset
)
```
Updates a LSTM gate for an iteration step of LSTM function, int16x8_16 version.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const int16_t * | in | Data input pointer |
| hidden_in | const int16_t * | in | Hidden state/ recurrent input pointer |
| gate_data | const cmsis_nn_lstm_gate * | in | Struct containing all information about the gate caluclation, see arm_nn_types. |
| params | const cmsis_nn_lstm_params * | in | Struct containing all information about the lstm_operation, see arm_nn_types |
| output | int16_t * | out | Hidden state/ recurrent output pointer |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches, see arm_nn_lstm_step_s16. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnsupportfunctions.h:3088`
## arm_nn_vec_mat_mul_result_acc_s8_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_vec_mat_mul_result_acc_s8_s16(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *effective_bias,
int16_t *dst,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t batches,
const int32_t batch_offset
)
```
The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a "batched" vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Batched vector |
| rhs | const int8_t * | in | Weights - input matrix (H(Rows)xW(Columns)) |
| effective_bias | const int32_t * | in | Bias + lhs_offset * kernel_sum term precalculated into a constant vector. |
| dst | int16_t * | out | Output |
| dst_multiplier | const int32_t | in | Multiplier for quantization |
| dst_shift | const int32_t | in | Shift for quantization |
| rhs_cols | const int32_t | in | Vector/matarix column length |
| rhs_rows | const int32_t | in | Row count of matrix |
| batches | const int32_t | in | Batch size |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches in input, see arm_nn_lstm_step_s8. Note that the output is always stored with sequential batches. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:3114`
## arm_nn_vec_mat_mul_result_acc_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_nn_vec_mat_mul_result_acc_s16(
const int16_t *lhs,
const int8_t *rhs,
const int64_t *effective_bias,
int16_t *dst,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t batches,
const int32_t batch_offset
)
```
The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a "batched" vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Batched vector |
| rhs | const int8_t * | in | Weights - input matrix (H(Rows)xW(Columns)) |
| effective_bias | const int64_t * | in | Bias + lhs_offset * kernel_sum term precalculated into a constant vector. |
| dst | int16_t * | out | Output |
| dst_multiplier | const int32_t | in | Multiplier for quantization |
| dst_shift | const int32_t | in | Shift for quantization |
| rhs_cols | const int32_t | in | Vector/matarix column length |
| rhs_rows | const int32_t | in | Row count of matrix |
| batches | const int32_t | in | Batch size |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches in input, see arm_nn_lstm_step_s16. Note that the output is always stored with sequential batches. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns `ARM_CMSIS_NN_SUCCESS` |
Source: `Include/arm_nnsupportfunctions.h:3144`
## arm_elementwise_mul_s16_s8
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_mul_s16_s8(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t block_size,
const int32_t batch_size,
const int32_t batch_offset
)
```
s16 elementwise multiplication with s8 output
Supported framework: TensorFlow Lite micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| output | int8_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| block_size | const int32_t | in | number of samples per batch |
| batch_size | const int32_t | in | number of samples per batch |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches in output, see arm_nn_lstm_step_s8. Note that it is assumed that the input is stored with sequential batches. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnsupportfunctions.h:3171`
## arm_elementwise_mul_s16_batch_offset
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_mul_s16_batch_offset(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t block_size,
const int32_t batch_size,
const int32_t batch_offset
)
```
s16 elementwise multiplication with s16 output
Supported framework: TensorFlow Lite micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| output | int16_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| block_size | const int32_t | in | number of samples per batch |
| batch_size | const int32_t | in | number of samples per batch |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches in output, see arm_nn_lstm_step_s16. Note that it is assumed that the input is stored with sequential batches. |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnsupportfunctions.h:3197`
## arm_elementwise_mul_acc_s16
`function` · `c`
```c
arm_cmsis_nn_status arm_elementwise_mul_acc_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)
```
s16 elementwise multiplication. The result of the multiplication is accumulated to the passed result buffer.
Supported framework: TensorFlow Lite micro
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1. Not used. |
| input_2_offset | const int32_t | in | offset for input 2. Not used. |
| output | int16_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Not used. |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns ARM_CMSIS_NN_SUCCESS |
Source: `Include/arm_nnsupportfunctions.h:3224`
## arm_check_broadcast_required
`function` · `c`
```c
static int32_t arm_check_broadcast_required(const cmsis_nn_dims *shape_1, const cmsis_nn_dims *shape_2)
```
Check if a broadcast is required between 2 `cmsis_nn_dims`.
Compares each dimension and returns 1 if any dimension does not match. This function does not check that broadcast rules are met.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| shape_1 | const cmsis_nn_dims * | in | pointer to input tensor 1 |
| shape_2 | const cmsis_nn_dims * | in | pointer to input tensor 2 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | The function returns 1 if a broadcast is required, or 0 if not. |
Source: `Include/arm_nnsupportfunctions.h:3245`
## arm_reduce_get_middle_block_from_arrays
`function` · `c`
```c
static int32_t arm_reduce_get_middle_block_from_arrays(
const int32_t in_dims,
const int32_t axis_arr,
int32_t *outer,
int32_t *reduce,
int32_t *inner
)
```
Reports whether the reduced axes of a 4-D tensor form one contiguous block followed by kept axes, as in a NHWC mean over H and W, and gives the flattened sizes. Axes of size 1 are ignored.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_dims | const int32_t | in | 4-element array {n, h, w, c} |
| axis_arr | const int32_t | in | 4-element mask {axis_n, axis_h, axis_w, axis_c} |
| outer | int32_t * | out | Product of the dims before the reduced block |
| reduce | int32_t * | out | Product of the reduced dims |
| inner | int32_t * | out | Product of the dims after the reduced block |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | 1 if the input is [outer, reduce, inner] with the middle dim reduced and inner > 1, otherwise 0 |
Source: `Include/arm_nnsupportfunctions.h:3310`
## ARM_NN_SQRT_S16_TABLEFREE_SHIFT
`macro` · `c`
```c
#define ARM_NN_SQRT_S16_TABLEFREE_SHIFT 14
```
Source: `Include/arm_nnsupportfunctions.h:3364`
## ARM_NN_SQRT_S16_TABLEFREE_MAGIC
`macro` · `c`
```c
#define ARM_NN_SQRT_S16_TABLEFREE_MAGIC UINT32_C(0x5F5FB6C4)
```
Source: `Include/arm_nnsupportfunctions.h:3365`
## ARM_NN_SQRT_S16_TABLEFREE_K0
`macro` · `c`
```c
#define ARM_NN_SQRT_S16_TABLEFREE_K0 (-4.76426697f)
```
Source: `Include/arm_nnsupportfunctions.h:3366`
## ARM_NN_SQRT_S16_TABLEFREE_K1
`macro` · `c`
```c
#define ARM_NN_SQRT_S16_TABLEFREE_K1 (-48.0000114f)
```
Source: `Include/arm_nnsupportfunctions.h:3367`
## arm_nn_sqrt_s16_tablefree_element
`function` · `c`
```c
static int16_t arm_nn_sqrt_s16_tablefree_element(const int32_t value, const float scale)
```
One element of `arm_sqrt_s16_tablefree()`: the float32 chain the MVE path evaluates per lane, so the two agree bit for bit on any IEEE-754 float32 implementation with round-to-nearest-even and a fused multiply-add (fmaf). Every product after the pre-scale either has two uses or feeds an fmaf or a conversion, never another lone multiply, so a compiler allowed to reassociate (-ffast-math) still has no chain to reorder, and no product feeds a bare add, so there is nothing to contract.
**Parameters**
| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| value | const int32_t | in | input code; values <= 0 give 0 |
| scale | const float | in | input_scale / (output_scale * output_scale) as float32 |
**Returns**
| Name | Type | Description |
| --- | --- | --- |
| | | trunc(sqrt(value * scale)) saturated to 32767 |
Source: `Include/arm_nnsupportfunctions.h:3381`