# heliaCORE.arm_nnsupportfunctions

## USE_FAST_DW_CONV_S16_FUNCTION

`macro` · `c`

```c
#define USE_FAST_DW_CONV_S16_FUNCTION(dw_conv_params, filter_dims, input_dims, output_dims) (dw_conv_params->ch_mult == 1 && \ arm_nn_dw_conv_opt_dilation_supported(dw_conv_params, input_dims, filter_dims, output_dims) && \ filter_dims->w * filter_dims->h < 512)
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params |  |  |  |
| filter_dims |  |  |  |
| input_dims |  |  |  |
| output_dims |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:45`

## LEFT_SHIFT

`macro` · `c`

```c
#define LEFT_SHIFT(_shift) (_shift > 0 ? _shift : 0)
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| _shift |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:50`

## RIGHT_SHIFT

`macro` · `c`

```c
#define RIGHT_SHIFT(_shift) (_shift > 0 ? 0 : -_shift)
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| _shift |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:51`

## MASK_IF_ZERO

`macro` · `c`

```c
#define MASK_IF_ZERO(x) (x) == 0 ? ~0 : 0
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:52`

## MASK_IF_NON_ZERO

`macro` · `c`

```c
#define MASK_IF_NON_ZERO(x) (x) != 0 ? ~0 : 0
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:53`

## SELECT_USING_MASK

`macro` · `c`

```c
#define SELECT_USING_MASK(mask, a, b) ((mask) & (a)) ^ (~(mask) & (b))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| mask |  |  |  |
| a |  |  |  |
| b |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:54`

## ARM_NN_MAX

`macro` · `c`

```c
#define ARM_NN_MAX(A, B) ((A) > (B) ? (A) : (B))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| A |  |  |  |
| B |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:57`

## ARM_NN_MIN

`macro` · `c`

```c
#define ARM_NN_MIN(A, B) ((A) < (B) ? (A) : (B))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| A |  |  |  |
| B |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:58`

## ARM_NN_CLAMP

`macro` · `c`

```c
#define ARM_NN_CLAMP(x, h, l) ARM_NN_MAX(ARM_NN_MIN((x), (h)), (l))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x |  |  |  |
| h |  |  |  |
| l |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:59`

## arm_nn_min_f16h

`function` · `c`

```c
static _Float16 arm_nn_min_f16h(_Float16 a, _Float16 b)
```

Minimum of two scalar f16 values.

With ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 minNum via VMINNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MIN, an ordered compare, so its NaN handling depends on operand order: a NaN `b` is returned, a NaN `a` is not. Do not rely on NaN suppression on non-MVE builds.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | _Float16 | in | First operand |
| b | _Float16 | in | Second operand |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The smaller of `a` and `b` |

Source: `Include/arm_nnsupportfunctions.h:98`

## arm_nn_max_f16h

`function` · `c`

```c
static _Float16 arm_nn_max_f16h(_Float16 a, _Float16 b)
```

Maximum of two scalar f16 values.

With ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 maxNum via VMAXNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MAX, an ordered compare, so its NaN handling depends on operand order: a NaN `b` is returned, a NaN `a` is not. Do not rely on NaN suppression on non-MVE builds.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | _Float16 | in | First operand |
| b | _Float16 | in | Second operand |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The larger of `a` and `b` |

Source: `Include/arm_nnsupportfunctions.h:121`

## arm_nn_propagate_nan_f16h

`function` · `c`

```c
static _Float16 arm_nn_propagate_nan_f16h(_Float16 x, _Float16 y)
```

Returns `x` when `x` is NaN, otherwise `y`.

Both the NaN test and the select are performed on the bit patterns: the test is (bits & 0x7FFF) > 0x7C00 (all-ones exponent, non-zero mantissa), which is integer arithmetic that -ffinite-math-only (implied by the shipped -Ofast) has no license to fold, unlike the former floating-point self-compare `x != x` (#333 / #334); the bit-pattern select neither expands to an HFmode conditional move (PR target/118460) nor quiets/retags the NaN payload. This helper backs the f16 elementwise clamp and, via arm_nn_clamp_scalar_f16 / arm_nn_clamp_propagate_nan_f16h, the other f16 scalar clamp users  arm_svdf_f16, arm_max_pool_f16 / arm_avg_pool_f16, the packed f16 matmul (arm_nn_mat_mult_nt_n_packed_f16), the scalar f16 RELU/RELU6/LEAKY_RELU activation legs, and arm_nn_vector_clamp_f16's scalar leg (conv/depthwise/transpose-conv f16, the 3x3 depthwise, and arm_nn_maxpool1d_f16)  so wherever that scalar clamp runs, a NaN passes through it at every optimization level on the gated toolchains. That is a guarantee about the clamp, not the whole kernel: which builds run the scalar clamp, and whether a NaN survives the rest of the kernel to reach it, is per kernel  several of these callers clamp with vmaxnmq/vminnmq on MVE builds (a NaN resolves to a bound there), and arm_max_pool_f16's max reduction drops a NaN before the clamp. The kernels with a NaN

:::note
(svdf, max/avg pool, packed matmul, activation) state their exact scope there; the arm_nn_vector_clamp_f16 family is covered by a test assertion in the transpose-conv f16 suite rather than per-kernel notes. The cortex-m55 MVE RELU/RELU6 f16 legs reach the same guarantee by the vector form of this idiom rather than by calling this helper: they restore NaN lanes with arm_nn_max_propagate_nan_mve_f16 / arm_nn_clamp_propagate_nan_mve_f16 (#382). The same idiom (bit-classified select) appears in arm_prelu_f16, which does not call this helper.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | _Float16 | in | Value whose NaN-ness selects the result. Returned unchanged when it is NaN. |
| y | _Float16 | in | Value returned when `x` is not NaN |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `x` if `x` is NaN, otherwise `y` |

Source: `Include/arm_nnsupportfunctions.h:167`

## arm_nn_clamp_f16h

`function` · `c`

```c
static _Float16 arm_nn_clamp_f16h(_Float16 x, _Float16 h, _Float16 l)
```

Drop-in equivalent of `ARM_NN_CLAMP(x, h, l)` for scalar _Float16 operands.

Includes the macro's NaN behaviour: `ARM_NN_MIN(NaN, h)` is h, so a NaN input resolves to the high bound, exactly as the macro does. Use arm_nn_clamp_propagate_nan_f16h() where TFLite NaN propagation is required.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | _Float16 | in | Value to clamp |
| h | _Float16 | in | Upper bound |
| l | _Float16 | in | Lower bound |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `x` clamped to [`l`, `h`] |

Source: `Include/arm_nnsupportfunctions.h:193`

## arm_nn_clamp_propagate_nan_f16h

`function` · `c`

```c
static _Float16 arm_nn_clamp_propagate_nan_f16h(_Float16 x, _Float16 l, _Float16 h)
```

Scalar f16 clamp with TFLite NaN semantics: NaN passes through unchanged.

Mirrors the MVE idiom in arm_nn_clamp_propagate_nan_mve_f16() (lower bound first, then upper bound, then restore NaN lanes). The NaN restore in arm_nn_propagate_nan_f16h() tests the integer bit pattern, so it holds at every optimization level including the shipped -Ofast; see #333 / #334. Bounds are assumed ordered (l <= h); inverted bounds are unspecified.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | _Float16 | in | Value to clamp |
| l | _Float16 | in | Lower bound |
| h | _Float16 | in | Upper bound |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `x` clamped to [`l`, `h`], or `x` itself when it is NaN |

Source: `Include/arm_nnsupportfunctions.h:212`

## arm_nn_abs_f16h

`function` · `c`

```c
static _Float16 arm_nn_abs_f16h(_Float16 x)
```

Absolute value of a scalar f16 value.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | _Float16 | in | Input value |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | \|`x`\| |

Source: `Include/arm_nnsupportfunctions.h:224`

## ARM_NN_ROUND_UP

`macro` · `c`

```c
#define ARM_NN_ROUND_UP(x, multiple) ((((x) + (multiple) - 1) / (multiple)) * (multiple))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x |  |  |  |
| multiple |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:236`

## REDUCE_MULTIPLIER

`macro` · `c`

```c
#define REDUCE_MULTIPLIER(_mult) ((_mult < 0x7FFF0000) ? ((_mult + (1 << 15)) >> 16) : 0x7FFF)
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| _mult |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:237`

## CH_IN_BLOCK_MVE

`macro` · `c`

```c
#define CH_IN_BLOCK_MVE (124)
```

Source: `Include/arm_nnsupportfunctions.h:245`

## S4_CH_IN_BLOCK_MVE

`macro` · `c`

```c
#define S4_CH_IN_BLOCK_MVE (124)
```

Source: `Include/arm_nnsupportfunctions.h:250`

## MAX_COL_COUNT

`macro` · `c`

```c
#define MAX_COL_COUNT (512)
```

Source: `Include/arm_nnsupportfunctions.h:254`

## REVERSE_TCOL_EFFICIENT_THRESHOLD

`macro` · `c`

```c
#define REVERSE_TCOL_EFFICIENT_THRESHOLD (16)
```

Source: `Include/arm_nnsupportfunctions.h:258`

## CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD

`macro` · `c`

```c
#define CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (1)
```

Source: `Include/arm_nnsupportfunctions.h:266`

## OPTIONAL_RESTRICT_KEYWORD

`macro` · `c`

```c
#define OPTIONAL_RESTRICT_KEYWORD
```

Source: `Include/arm_nnsupportfunctions.h:272`

## arm_nn_size_mul

`function` · `c`

```c
static int64_t arm_nn_size_mul(const int64_t acc, const int64_t factor)
```

Fold one dimension into a running buffer-size product, reporting overflow as -1.

Buffer-size queries return an int32_t byte count, so the product of the dimensions they multiply has to be rejected as soon as it cannot fit. Folding one factor at a time keeps the accumulator bounded: an accumulator already known to be <= INT32_MAX times a factor <= INT32_MAX cannot exceed about 2^62, so the int64_t accumulator itself never wraps. Chaining raw (int64_t) casts across three or more int32_t dims does not have that property - 65536 * 65536 * 65536 * 65536 is exactly 2^64 and folds back to 0, which would sail through a trailing "> INT32_MAX" test.

:::note
This is the -1 sentinel family, used by the s8/s16 integer buffer-size queries, by the eight SVDF staging queries (arm_svdf_{s8,state_s16_s8,f32,f16}_{input,output}_ctx_get_buffer_size) and by the s8/s16 LSTM temp-buffer queries and the GRU temp queries (arm_lstm_unidirectional_{s8,s16}_temp{1,2}_get_buffer_size, arm_gru_unidirectional_{f32,f16}_temp1_get_buffer_size). The four f32/f16 LSTM temp queries have no dimensions to fold (the buffers are unused) and answer -1 only for NULL params, 0 otherwise. It is not interchangeable with the arm_nn_checked_size_mul() / arm_nn_size_to_i32_or_zero() helpers in Source/NNSupportFunctions (shared header for the float sizers), which most f32 and f16 buffer-size queries use and which report an out-of-range size as 0. Mixing the two silently flips a sizer's out-of-range contract from "must never be used to size a buffer" to "you may pass { NULL, 0 }", so pick the one the surrounding family already uses.

:::

:::note
The split is per sizer, not per datatype. The four SVDF f32/f16 staging queries deliberately use this -1 family rather than the 0 one their neighbours use, because their kernels read ctx->size and a size of 0 opts out of the scratch-size check - so a 0-on-overflow answer fed back as { alloc(0), 0 } would disable the very check meant to catch it. Do not infer a sizer's sentinel from its datatype suffix.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| acc | const int64_t | in | Running product, or -1 if an earlier fold already overflowed. |
| factor | const int64_t | in | Next factor to fold in. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | acc * factor, or -1 if acc is already -1, factor is negative or out of int32_t range, or the product exceeds INT32_MAX. |

Source: `Include/arm_nnsupportfunctions.h:307`

## arm_nn_size_add

`function` · `c`

```c
static int64_t arm_nn_size_add(const int64_t acc, const int64_t addend)
```

Add to a running buffer-size product, reporting overflow as -1.

Companion to arm_nn_size_mul() for the sizers that append a fixed slack term.

:::note
Same sentinel caveat as arm_nn_size_mul() - see its note for the full split, including why the four SVDF f32/f16 staging queries use this -1 family rather than the 0-returning arm_nn_checked_size_mul() / arm_nn_size_to_i32_or_zero() family that most other float sizers use.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| acc | const int64_t | in | Running product, or -1 if an earlier step already overflowed. |
| addend | const int64_t | in | Value to add. Must be non-negative. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | acc + addend, or -1 if acc is already -1 or the sum exceeds INT32_MAX. |

Source: `Include/arm_nnsupportfunctions.h:332`

## PACK_S8x4_32x1

`macro` · `c`

```c
#define PACK_S8x4_32x1(v0, v1, v2, v3) ((int32_t)((((uint32_t)(v0)) & 0xFFu) | ((((uint32_t)(v1)) & 0xFFu) << 8) | ((((uint32_t)(v2)) & 0xFFu) << 16) | \ ((((uint32_t)(v3)) & 0xFFu) << 24)))
```

definition to pack four 8 bit values.

Byte lanes are masked and shifted in uint32_t so a negative value never feeds a signed left shift (UB); masking before the shift keeps the same bits the old shift-then-mask form kept. Bit-identical for every input. Deliberate divergence from upstream ARM-software/CMSIS-NN, which still carries the signed-shift form  do not paste the upstream text back on a sync (issue #357).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| v0 |  |  |  |
| v1 |  |  |  |
| v2 |  |  |  |
| v3 |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:358`

## PACK_Q15x2_32x1

`macro` · `c`

```c
#define PACK_Q15x2_32x1(v0, v1) ((int32_t)((((uint32_t)(v0)) & 0xFFFFu) | (((uint32_t)(v1)) << 16)))
```

definition to pack two 16 bit values.

Same treatment: the high half is shifted in uint32_t, not int32_t, so a negative v1 is defined; the low half keeps its mask. Bit-identical for every input. Same deliberate upstream divergence as PACK_S8x4_32x1 above.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| v0 |  |  |  |
| v1 |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:369`

## GetNearestNeighbor

`function` · `c`

```c
static int32_t GetNearestNeighbor(
    const int input_value,
    const int32_t input_size,
    const float scale,
    const float offset,
    const bool align_corners,
    const bool half_pixel_centers
)
```

Map an output index to the nearest input index for resize.

This helper follows the TensorFlow Lite nearest-neighbor resize mapping rules.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_value | const int | in | Output index (x or y). |
| input_size | const int32_t | in | Input size along the same axis. |
| scale | const float | in | Precomputed scaling factor for the axis. |
| offset | const float | in | Precomputed offset for the axis. |
| align_corners | const bool | in | If true, use align-corners scaling. |
| half_pixel_centers | const bool | in | If true, use half-pixel center offset. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Nearest input index for the given output index. |

Source: `Include/arm_nnsupportfunctions.h:391`

## arm_nn_is_convolve_1x1

`function` · `c`

```c
static bool arm_nn_is_convolve_1x1(
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims
)
```

Check if convolution parameters correspond to a 1x1 convolution.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input dimensions |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | true if parameters describe a 1x1 convolution, false otherwise. |

Source: `Include/arm_nnsupportfunctions.h:416`

## arm_nn_is_convolve_1x1_fast

`function` · `c`

```c
static bool arm_nn_is_convolve_1x1_fast(const cmsis_nn_conv_params *conv_params)
```

Check if a 1x1 convolution qualifies for the fast (unit stride) path.

:::note
Does not validate that the kernel is 1x1. Call arm_nn_is_convolve_1x1() first.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | true if stride is 1x1, false otherwise. |

Source: `Include/arm_nnsupportfunctions.h:432`

## arm_nn_is_convolve_1_x_n

`function` · `c`

```c
static bool arm_nn_is_convolve_1_x_n(
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims
)
```

Check if convolution parameters correspond to a 1xN convolution.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input dimensions |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | true if parameters describe a 1xN convolution, false otherwise. |

Source: `Include/arm_nnsupportfunctions.h:444`

## arm_nn_convolve_1_x_n_padding_supported

`function` · `c`

```c
static bool arm_nn_convolve_1_x_n_padding_supported(
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims
)
```

Check that `arm_convolve_1_x_n_s4()` handles the horizontal padding of a 1xN convolution.

The kernel places pad.w columns on the left and pad.w + (total_pad % 2) on the right, where total_pad = (output W - 1) * stride.w + filter W - input W, and needs the output columns that read padding to fit in output W. Its padded-column code also assumes that each such column reads at least one input column and that the filter is no wider than the input; otherwise it forms input and filter addresses outside the tensors. On MVE builds another pad placement, or too many padded columns, returns ARM_CMSIS_NN_FAILURE. A VALID layer whose stride leaves trailing input unused (negative total_pad) is therefore rejected, and the wrapper routes it to another convolution. The kernel also computes a single output row, so vertical padding or an output height other than 1 is rejected. A non-positive stride.w is left to the kernel's argument checks.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input dimensions |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output dimensions |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | true when `arm_convolve_1_x_n_s4()` handles the padding, false otherwise. |

Source: `Include/arm_nnsupportfunctions.h:473`

## arm_nn_convolve_1_x_n_s8_padding_supported

`function` · `c`

```c
static bool arm_nn_convolve_1_x_n_s8_padding_supported(
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims
)
```

Check that `arm_convolve_1_x_n_s8()` accepts the padding and output shape of a 1xN convolution.

The kernel computes a single output row for any pad.w >= 0 and any output width, including an odd total padding, a filter wider than the input and a VALID layer whose stride leaves trailing input unused. It rejects vertical padding, an output height other than 1, a negative pad.w and an empty filter; the wrapper routes those layers to another convolution. A non-positive stride.w is left to the kernel's argument checks.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output dimensions |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | true when `arm_convolve_1_x_n_s8()` computes the layer, false otherwise. |

Source: `Include/arm_nnsupportfunctions.h:517`

## arm_nn_convolve_1_x_n_padded_columns

`function` · `c`

```c
static void arm_nn_convolve_1_x_n_padded_columns(
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    int64_t *left_num,
    int64_t *right_num
)
```

Count the output columns of a 1xN convolution whose window reads padding.

Output column j reads input columns j * stride.w - pad.w to j * stride.w - pad.w + filter W - 1. The leading columns whose window starts before the input are left-padded; of the others, the trailing columns whose window ends past the input are right-padded. A window can do both only when it is left-padded.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters. stride.w >= 1 and pad.w >= 0. |
| input_dims | const cmsis_nn_dims * | in | Input dimensions. w >= 0. |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions. w >= 1. |
| output_dims | const cmsis_nn_dims * | in | Output dimensions. w >= 0. |
| left_num | int64_t * | out | Number of left-padded output columns, output W at most. |
| right_num | int64_t * | out | Number of right-padded output columns, output W - left_num at most. |

Source: `Include/arm_nnsupportfunctions.h:539`

## arm_nn_dw_conv_opt_dilation_supported

`function` · `c`

```c
static bool arm_nn_dw_conv_opt_dilation_supported(
    const cmsis_nn_dw_conv_params *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims
)
```

Check if the dilation, stride and padding of a depthwise layer allow the `arm_depthwise_conv_s8_opt()` or `arm_depthwise_conv_fast_s16()` route.

:::note
Does not check ch_mult, the batch count or the kernel size: `arm_depthwise_conv_wrapper_s8()`, `arm_depthwise_conv_wrapper_s16()` and their buffer-size functions apply their own conditions on those, and all of them take this predicate so that routing and sizing agree.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input dimensions |
| filter_dims | const cmsis_nn_dims * | in | Filter dimensions |
| output_dims | const cmsis_nn_dims * | in | Output dimensions |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | true for an undilated layer (dilation 1 in both dimensions), or for a 1D layer dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, no vertical padding, dilation.h == 1 and dilation.w >= 1. false otherwise. |

Source: `Include/arm_nnsupportfunctions.h:572`

## arm_q7_to_q15_with_offset

`function` · `c`

```c
void arm_q7_to_q15_with_offset(const int8_t *src, int16_t *dst, int32_t block_size, int16_t offset)
```

Converts the elements from a s8 vector to a s16 vector with an added offset.

Output elements are ordered. The equation used for the conversion process is:

dst[n] = (int16_t) src[n] + offset; 0 <= n < block_size.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| src | const int8_t * | in | pointer to the s8 input vector |
| dst | int16_t * | out | pointer to the s16 output vector |
| block_size | int32_t | in | length of the input vector |
| offset | int16_t | in | s16 offset to be added to each input vector element. |

Source: `Include/arm_nnsupportfunctions.h:649`

## arm_depthwise_conv_s8_opt_get_buffer_size_mve

`function` · `c`

```c
int32_t arm_depthwise_conv_s8_opt_get_buffer_size_mve(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```

Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with MVE extension.

The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.

:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_s8_opt_get_buffer_size()`. Note also this is a support function, so not recommended to call directly even on Host.

:::

:::note
This leg sizes its buffer from a fixed channel block rather than from input_dims->c, but it applies the same dimension check as `arm_depthwise_conv_s8_opt_get_buffer_size()` anyway, returning -1 for a negative input_dims->c or filter dimension, or for a byte count that would not fit in an int32_t. Without that check this entry point - and every s4 depthwise sizer, which route here - answered a negative channel count with a plausible positive size (issue #318).

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |

Source: `Include/arm_nnsupportfunctions.h:695`

## arm_depthwise_conv_s8_opt_get_buffer_size_dsp

`function` · `c`

```c
int32_t arm_depthwise_conv_s8_opt_get_buffer_size_dsp(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)
```

Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with DSP extension.

The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.

:::note
Intended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_s8_opt_get_buffer_size()`. Note also this is a support function, so not recommended to call directly even on Host.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |

Source: `Include/arm_nnsupportfunctions.h:710`

## arm_nn_depthwise_conv_s8_core

`function` · `c`

```c
int8_t * arm_nn_depthwise_conv_s8_core(
    const int8_t *row,
    const int16_t *col,
    const uint16_t num_ch,
    const int32_t *out_shift,
    const int32_t *out_mult,
    const int32_t out_offset,
    const int32_t activation_min,
    const int32_t activation_max,
    const uint16_t kernel_size,
    const int32_t *const output_bias,
    int8_t *out
)
```

Depthwise conv on an im2col buffer where the input channel equals output channel.

Supported framework: TensorFlow Lite micro.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| row | const int8_t * | in | pointer to row |
| col | const int16_t * | in | pointer to im2col buffer, always consists of 2 columns. |
| num_ch | const uint16_t | in | number of channels |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| activation_min | const int32_t | in | minimum value to clamp the output to. Range : int8 |
| activation_max | const int32_t | in | maximum value to clamp the output to. Range : int8 |
| kernel_size | const uint16_t | in | number of elements in one column. |
| output_bias | const int32_t *const | in | per output channel bias. Range : int32 |
| out | int8_t * | out | pointer to output |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns one of the two<br><br>1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |

Source: `Include/arm_nnsupportfunctions.h:732`

## arm_nn_mat_mult_s8

`function` · `c`

```c
int8_t * arm_nn_mat_mult_s8(
    const int8_t *input_row,
    const int8_t *input_col,
    const uint16_t output_ch,
    const uint16_t col_batches,
    const int32_t *output_shift,
    const int32_t *output_mult,
    const int32_t out_offset,
    const int32_t col_offset,
    const int32_t row_offset,
    const int16_t out_activation_min,
    const int16_t out_activation_max,
    const uint16_t row_len,
    const int32_t *const bias,
    int8_t *out
)
```

General Matrix-multiplication function with per-channel requantization.

Supported framework: TensorFlow Lite

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_row | const int8_t * | in | pointer to row operand |
| input_col | const int8_t * | in | pointer to col operand |
| output_ch | const uint16_t | in | number of rows of input_row |
| col_batches | const uint16_t | in | number of column batches. Range: 1 to 4 |
| output_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| output_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| col_offset | const int32_t | in | input tensor(col) offset. |
| row_offset | const int32_t | in | kernel offset(row). Not used. |
| out_activation_min | const int16_t | in | minimum value to clamp the output to. Range : int8 |
| out_activation_max | const int16_t | in | maximum value to clamp the output to. Range : int8 |
| row_len | const uint16_t | in | number of elements in each row |
| bias | const int32_t *const | in | per output channel bias. Range : int32 |
| out | int8_t * | in, out | pointer to output |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns one of the two<br><br>1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |

Source: `Include/arm_nnsupportfunctions.h:766`

## arm_nn_mat_mult_kernel_s16

`function` · `c`

```c
int16_t * arm_nn_mat_mult_kernel_s16(
    const int8_t *input_a,
    const int16_t *input_b,
    const int32_t output_ch,
    const int32_t *out_shift,
    const int32_t *out_mult,
    const int32_t activation_min,
    const int32_t activation_max,
    const int32_t num_col_a,
    const cmsis_nn_bias_data *const bias_data,
    int16_t *out_0,
    const int32_t row_address_offset
)
```

Matrix-multiplication function for convolution with per-channel requantization for 16 bits convolution.

This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_a | const int8_t * | in | pointer to operand A |
| input_b | const int16_t * | in | pointer to operand B, always consists of 2 vectors. |
| output_ch | const int32_t | in | number of rows of A |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| activation_min | const int32_t | in | minimum value to clamp the output to. Range : int16 |
| activation_max | const int32_t | in | maximum value to clamp the output to. Range : int16 |
| num_col_a | const int32_t | in | number of columns of A |
| bias_data | const cmsis_nn_bias_data *const | in | pointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct. |
| out_0 | int16_t * | in, out | pointer to output |
| row_address_offset | const int32_t | in | Address offset between rows in output. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns one of the two<br><br>1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |

Source: `Include/arm_nnsupportfunctions.h:804`

## arm_nn_mat_mul_core_1x_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mul_core_1x_s8(
    int32_t row_elements,
    const int32_t skipped_row_elements,
    const int8_t *row_base_ref,
    const int8_t *col_base_ref,
    const int32_t out_ch,
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_per_channel_quant_params *quant_params,
    const int32_t *bias,
    int8_t *output
)
```

General Vector by Matrix multiplication with requantization and storage of result.

Pseudo-code *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in 'output'

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| row_elements | int32_t | in | number of row elements |
| skipped_row_elements | const int32_t | in | number of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch |
| row_base_ref | const int8_t * | in | pointer to row operand |
| col_base_ref | const int8_t * | in | pointer to col operand |
| out_ch | const int32_t | in | Number of output channels |
| conv_params | const cmsis_nn_conv_params * | in | Pointer to convolution parameters like offsets and activation values |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Pointer to per-channel quantization parameters |
| bias | const int32_t * | in | Pointer to optional per-channel bias |
| output | int8_t * | out | Pointer to output where int8 results are stored. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory. |

Source: `Include/arm_nnsupportfunctions.h:843`

## arm_nn_mat_mul_core_1x_s4

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mul_core_1x_s4(
    int32_t row_elements,
    const int32_t skipped_row_elements,
    const int8_t *row_base_ref,
    const int8_t *col_base_ref,
    const int32_t out_ch,
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_per_channel_quant_params *quant_params,
    const int32_t *bias,
    int8_t *output
)
```

General Vector by Matrix multiplication with requantization, storage of result and int4 weights packed into an int8 buffer.

Pseudo-code as int8 example. Int4 filter data will be unpacked. *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in 'output'

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| row_elements | int32_t | in | number of row elements |
| skipped_row_elements | const int32_t | in | number of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch |
| row_base_ref | const int8_t * | in | pointer to row operand |
| col_base_ref | const int8_t * | in | pointer to col operand as packed int4 |
| out_ch | const int32_t | in | Number of output channels |
| conv_params | const cmsis_nn_conv_params * | in | Pointer to convolution parameters like offsets and activation values |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Pointer to per-channel quantization parameters |
| bias | const int32_t * | in | Pointer to optional per-channel bias |
| output | int8_t * | out | Pointer to output where int8 results are stored. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory. |

Source: `Include/arm_nnsupportfunctions.h:881`

## arm_nn_mat_mul_core_4x_s8

`function` · `c`

```c
int8_t * arm_nn_mat_mul_core_4x_s8(
    const int32_t row_elements,
    const int32_t offset,
    const int8_t *row_base,
    const int8_t *col_base,
    const int32_t out_ch,
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_per_channel_quant_params *quant_params,
    const int32_t *bias,
    int8_t *output
)
```

Matrix-multiplication with requantization & activation function for four rows and one column.

Compliant to TFLM int8 specification. MVE implementation only

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| row_elements | const int32_t | in | number of row elements |
| offset | const int32_t | in | offset between rows. Can be the same as row_elements. For e.g, in a 1x1 conv scenario with stride as 1. |
| row_base | const int8_t * | in | pointer to row operand |
| col_base | const int8_t * | in | pointer to col operand |
| out_ch | const int32_t | in | Number of output channels |
| conv_params | const cmsis_nn_conv_params * | in | Pointer to convolution parameters like offsets and activation values |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Pointer to per-channel quantization parameters |
| bias | const int32_t * | in | Pointer to per-channel bias |
| output | int8_t * | out | Pointer to output where int8 results are stored. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns the updated output pointer or NULL if implementation is not available. |

Source: `Include/arm_nnsupportfunctions.h:908`

## arm_nn_mat_mult_nt_t_s4

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s4(
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t *bias,
    int8_t *dst,
    const int32_t *dst_multipliers,
    const int32_t *dst_shifts,
    const int32_t lhs_rows,
    const int32_t rhs_rows,
    const int32_t rhs_cols,
    const int32_t lhs_offset,
    const int32_t dst_offset,
    const int32_t activation_min,
    const int32_t activation_max,
    const int32_t lhs_cols_offset
)
```

General Matrix-multiplication function with per-channel requantization. This function assumes:

- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)
- RHS is int8 packed with 2x int4
- LHS is int8

:::note
This operation also performs the broadcast bias addition before the requantization

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias | const int32_t * | in | Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst | int8_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_offset | const int32_t | in | Offset to be applied the output result |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int8 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int8 |
| lhs_cols_offset | const int32_t | in | Column offset between subsequent lhs_rows |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:950`

## arm_nn_mat_mult_nt_interleaved_t_even_s4

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_interleaved_t_even_s4(
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t *bias,
    int8_t *dst,
    const int32_t *dst_multipliers,
    const int32_t *dst_shifts,
    const int32_t lhs_rows,
    const int32_t rhs_rows,
    const int32_t rhs_cols,
    const int32_t lhs_offset,
    const int32_t dst_offset,
    const int32_t activation_min,
    const int32_t activation_max,
    const int32_t lhs_cols_offset
)
```

General Matrix-multiplication function with per-channel requantization. This function assumes:

- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)
- RHS is int8 packed with 2x int4
- LHS is int8
- LHS/RHS input columns must be even numbered
- LHS must be interleaved. Compare to arm_nn_mat_mult_nt_t_s4 where LHS is not interleaved.

:::note
This operation also performs the broadcast bias addition before the requantization

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias | const int32_t * | in | Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst | int8_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns. Note this must be even. |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_offset | const int32_t | in | Offset to be applied the output result |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int8 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int8 |
| lhs_cols_offset | const int32_t | in | Column offset between subsequent lhs_rows |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:999`

## arm_nn_mat_mult_nt_t_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s8(
    const int32_t *weight_sum_buf,
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t *bias,
    int8_t *dst,
    const int32_t *dst_multipliers,
    const int32_t *dst_shifts,
    const int32_t lhs_rows,
    const int32_t rhs_rows,
    const int32_t rhs_cols,
    const int32_t lhs_offset,
    const int32_t dst_offset,
    const int32_t activation_min,
    const int32_t activation_max,
    const int32_t row_address_offset,
    const int32_t lhs_cols_offset
)
```

General Matrix-multiplication function with per-channel requantization. This function assumes:

- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)

:::note
This operation also performs the broadcast bias addition before the requantization

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| weight_sum_buf | const int32_t * | in | Pointer to the weight sum multiplied by lhs_offset and summed bias buffer |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias | const int32_t * | in | Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst | int8_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_offset | const int32_t | in | Offset to be applied the output result |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int8 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int8 |
| row_address_offset | const int32_t | in | Address offset between rows in output. NOTE: Only used for MVEI extension. |
| lhs_cols_offset | const int32_t | in | Column offset between subsequent lhs_rows |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1046`

## arm_nn_mat_mult_nt_t_1x1_out_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_1x1_out_s8(
    const int32_t *weight_sum_buf,
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t *bias,
    int8_t *dst,
    const int32_t *dst_multipliers,
    const int32_t *dst_shifts,
    const int32_t lhs_rows,
    const int32_t rhs_rows,
    const int32_t rhs_cols,
    const int32_t lhs_offset,
    const int32_t dst_offset,
    const int32_t activation_min,
    const int32_t activation_max,
    const int32_t row_address_offset,
    const int32_t lhs_cols_offset
)
```

General Matrix-multiplication function with per-channel requantization. Output is calculated with multiple channels in parallel, rather than multiple output indices in a single channel This function assumes:

- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)

:::note
This operation also performs the broadcast bias addition before the requantization

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| weight_sum_buf | const int32_t * | in | Pointer to the weight sum multiplied by lhs_offset and summed bias buffer |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias | const int32_t * | in | Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst | int8_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_offset | const int32_t | in | Offset to be applied the output result |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int8 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int8 |
| row_address_offset | const int32_t | in | Address offset between rows in output. NOTE: Only used for MVEI extension. |
| lhs_cols_offset | const int32_t | in | Column offset between subsequent lhs_rows |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1097`

## arm_nn_mat_mult_nt_t_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s16(
    const int16_t *lhs,
    const int8_t *rhs,
    const cmsis_nn_bias_data *bias_data,
    int16_t *dst,
    const int32_t *dst_multipliers,
    const int32_t *dst_shifts,
    const int32_t lhs_rows,
    const int32_t rhs_rows,
    const int32_t rhs_cols,
    const int32_t activation_min,
    const int32_t activation_max,
    const int32_t row_address_offset
)
```

General Matrix-multiplication function with per-channel requantization and int16 input (LHS) and output. This function assumes:

- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)

:::note
This operation also performs the broadcast bias addition before the requantization

:::

MVE implementation only.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| bias_data | const cmsis_nn_bias_data * | in | Pointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct. |
| dst | int16_t * | out | Pointer to the output matrix with "m" rows and "n" columns |
| dst_multipliers | const int32_t * | in | Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| dst_shifts | const int32_t * | in | Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows) |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of RHS input rows |
| rhs_cols | const int32_t | in | Number of LHS/RHS input columns |
| activation_min | const int32_t | in | Minimum value to clamp down the output. Range : int16 |
| activation_max | const int32_t | in | Maximum value to clamp up the output. Range : int16 |
| row_address_offset | const int32_t | in | Address offset between rows in output. NOTE: Only used for MVEI extension. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` or `ARM_CMSIS_NN_NO_IMPL_ERROR` if not for MVE \|---row_address_offset---\| \|____rhs_rows__________________\|<br><br>\|  \|  \| \| --- \| --- \| \|  \|  \| \|  \|  \|<br><br>\| \| \| lhs_rows<br><br>\|  \|  \| \| --- \| --- \| \| _______________ \| ______________ \| |

Source: `Include/arm_nnsupportfunctions.h:1155`

## arm_nn_mat_mult_nt_t_s8_s32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s8_s32(
    const int8_t *lhs,
    const int8_t *rhs,
    int32_t *dst,
    const int32_t lhs_rows,
    const int32_t rhs_rows,
    const int32_t rhs_cols,
    const int32_t lhs_offset,
    const int32_t dst_idx_offset
)
```

General Matrix-multiplication function with int8 input and int32 output. This function assumes:

- LHS input matrix NOT transposed (nt)
- RHS input matrix transposed (t)

:::note
Dst/output buffer must be zeroed out before calling this function.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Pointer to the LHS input matrix |
| rhs | const int8_t * | in | Pointer to the RHS input matrix |
| dst | int32_t * | in, out | Pointer to the output matrix with "m" rows and "n" columns. Accumulated into, so it must be zeroed by the caller before the call |
| lhs_rows | const int32_t | in | Number of LHS input rows |
| rhs_rows | const int32_t | in | Number of LHS input columns/RHS input rows |
| rhs_cols | const int32_t | in | Number of RHS input columns |
| lhs_offset | const int32_t | in | Offset to be applied to the LHS input value |
| dst_idx_offset | const int32_t | in | Offset between subsequent output results |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1189`

## arm_nn_vec_mat_mult_t_s4

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s4(
    const int8_t *lhs,
    const int8_t *packed_rhs,
    const int32_t *bias,
    int8_t *dst,
    const int32_t lhs_offset,
    const int32_t dst_offset,
    const int32_t dst_multiplier,
    const int32_t dst_shift,
    const int32_t rhs_cols,
    const int32_t rhs_rows,
    const int32_t activation_min,
    const int32_t activation_max
)
```

s4 Vector by Matrix (transposed) multiplication

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side vector |
| packed_rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| bias | const int32_t * | in | Input bias |
| dst | int8_t * | out | Output vector |
| lhs_offset | const int32_t | in | Offset to be added to the input values of the left-hand side vector. Range: -127 to 128 |
| dst_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1218`

## arm_nn_vec_mat_mult_t_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s8(
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t *kernel_sum,
    const int32_t *bias,
    int8_t *dst,
    const int32_t lhs_offset,
    const int32_t dst_offset,
    const int32_t dst_multiplier,
    const int32_t dst_shift,
    const int32_t rhs_cols,
    const int32_t rhs_rows,
    const int32_t activation_min,
    const int32_t activation_max,
    const int32_t address_offset,
    const int32_t rhs_offset
)
```

s8 Vector by Matrix (transposed) multiplication

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| kernel_sum | const int32_t * | in | Kernel sums of the kernels (rhs). See arm_vector_sum_s8 for more info. |
| bias | const int32_t * | in | Input bias |
| dst | int8_t * | out | Output vector |
| lhs_offset | const int32_t | in | Offset to be added to the input values of the left-hand side vector. Range: -127 to 128 |
| dst_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| address_offset | const int32_t | in | Memory position offset for dst. First output is stored at 'dst', the second at 'dst + address_offset' and so on. Default value is typically 1. |
| rhs_offset | const int32_t | in | Offset to be added to the input values of the right-hand side vector. Range: -127 to 128 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1256`

## arm_nn_vec_mat_mult_t_per_ch_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_per_ch_s8(
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t *kernel_sum,
    const int32_t *bias,
    int8_t *dst,
    const int32_t lhs_offset,
    const int32_t dst_offset,
    const int32_t *dst_multiplier,
    const int32_t *dst_shift,
    const int32_t rhs_cols,
    const int32_t rhs_rows,
    const int32_t activation_min,
    const int32_t activation_max,
    const int32_t address_offset,
    const int32_t rhs_offset
)
```

s8 Vector by Matrix (transposed) multiplication using per channel quantization for output

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| kernel_sum | const int32_t * | in | Kernel sums of the kernels (rhs). See arm_vector_sum_s8 for more info. |
| bias | const int32_t * | in | Input bias |
| dst | int8_t * | out | Output vector |
| lhs_offset | const int32_t | in | Offset to be added to the input values of the left-hand side vector. Range: -127 to 128 |
| dst_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| dst_multiplier | const int32_t * | in | Output multipliers |
| dst_shift | const int32_t * | in | Output shifts |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| address_offset | const int32_t | in | Memory position offset for dst. First output is stored at 'dst', the second at 'dst + address_offset' and so on. Default value is typically 1. |
| rhs_offset | const int32_t | in | Offset to be added to the input values of the right-hand side vector. Range: -127 to 128 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1297`

## arm_nn_vec_mat_mult_t_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s16(
    const int16_t *lhs,
    const int8_t *rhs,
    const int64_t *bias,
    int16_t *dst,
    const int32_t dst_multiplier,
    const int32_t dst_shift,
    const int32_t rhs_cols,
    const int32_t rhs_rows,
    const int32_t activation_min,
    const int32_t activation_max
)
```

s16 Vector by s8 Matrix (transposed) multiplication

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| bias | const int64_t * | in | Input bias |
| dst | int16_t * | out | Output vector |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int16 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int16 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1330`

## arm_nn_vec_mat_mult_t_per_ch_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_per_ch_s16(
    const int16_t *lhs,
    const int8_t *rhs,
    const int64_t *bias,
    int16_t *dst,
    const int32_t *dst_multiplier,
    const int32_t *dst_shift,
    const int32_t rhs_cols,
    const int32_t rhs_rows,
    const int32_t activation_min,
    const int32_t activation_max
)
```

s16 vector(lhs) by s8 matrix (transposed) multiplication and per channel quant output

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| bias | const int64_t * | in | Input bias |
| dst | int16_t * | out | Output vector |
| dst_multiplier | const int32_t * | in | Per channel output multiplier. Length of vector is equal to rhs_rows |
| dst_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to rhs_rows |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int16 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int16 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1358`

## arm_nn_vec_mat_mult_t_s16_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s16_s16(
    const int16_t *lhs,
    const int16_t *rhs,
    const int64_t *bias,
    int16_t *dst,
    const int32_t dst_multiplier,
    const int32_t dst_shift,
    const int32_t rhs_cols,
    const int32_t rhs_rows,
    const int32_t activation_min,
    const int32_t activation_max
)
```

s16 Vector by s16 Matrix (transposed) multiplication

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Input left-hand side vector |
| rhs | const int16_t * | in | Input right-hand side matrix (transposed) |
| bias | const int64_t * | in | Input bias |
| dst | int16_t * | out | Output vector |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int16 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int16 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1386`

## arm_nn_vec_mat_mult_t_svdf_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_vec_mat_mult_t_svdf_s8(
    const int8_t *lhs,
    const int8_t *rhs,
    int16_t *dst,
    const int32_t lhs_offset,
    const int32_t scatter_offset,
    const int32_t dst_multiplier,
    const int32_t dst_shift,
    const int32_t rhs_cols,
    const int32_t rhs_rows,
    const int32_t activation_min,
    const int32_t activation_max
)
```

s8 Vector by Matrix (transposed) multiplication with s16 output

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side vector |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| dst | int16_t * | out | Output vector |
| lhs_offset | const int32_t | in | Offset to be added to the input values of the left-hand side vector. Range: -127 to 128 |
| scatter_offset | const int32_t | in | Address offset for dst. First output is stored at 'dst', the second at 'dst + scatter_offset' and so on. |
| dst_multiplier | const int32_t | in | Output multiplier |
| dst_shift | const int32_t | in | Output shift |
| rhs_cols | const int32_t | in | Number of columns in the right-hand side input matrix |
| rhs_rows | const int32_t | in | Number of rows in the right-hand side input matrix |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int16 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int16 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:1417`

## arm_nn_depthwise_conv_nt_t_padded_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_padded_s8(
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t lhs_offset,
    const int32_t active_ch,
    const int32_t total_ch,
    const int32_t *out_shift,
    const int32_t *out_mult,
    const int32_t out_offset,
    const int32_t activation_min,
    const int32_t activation_max,
    const uint16_t row_x_col,
    const int32_t *const output_bias,
    int8_t *out
)
```

Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in padded cases where the padding is -lhs_offset(Range: int8). Dimensions are the same for lhs and rhs.

:::note
Tail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `active_ch`.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side matrix |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| lhs_offset | const int32_t | in | LHS matrix offset(input offset). Range: -127 to 128 |
| active_ch | const int32_t | in | Subset of total_ch processed |
| total_ch | const int32_t | in | Number of channels in LHS/RHS |
| out_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to number of channels |
| out_mult | const int32_t * | in | Per channel output multiplier. Length of vector is equal to number of channels |
| out_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| row_x_col | const uint16_t | in | (row_dimension * col_dimension) of LHS/RHS matrix |
| output_bias | const int32_t *const | in | Per channel output bias. Length of vector is equal to number of channels |
| out | int8_t * | out | Output pointer |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` if an implementation is available or `ARM_CMSIS_NN_NO_IMPL_ERROR` otherwise |

Source: `Include/arm_nnsupportfunctions.h:1453`

## arm_nn_depthwise_conv_nt_t_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_s8(
    const int32_t *weight_sum_buf,
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t lhs_offset,
    const int32_t active_ch,
    const int32_t total_ch,
    const int32_t *out_shift,
    const int32_t *out_mult,
    const int32_t out_offset,
    const int32_t activation_min,
    const int32_t activation_max,
    const uint16_t row_x_col,
    const int32_t *const output_bias,
    int8_t *out
)
```

Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. Dimensions are the same for lhs and rhs.

:::note
Tail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `active_ch`.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| weight_sum_buf | const int32_t * | in | Pointer to the weight sum multiplied by lhs_offset and summed bias buffer |
| lhs | const int8_t * | in | Input left-hand side matrix |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| lhs_offset | const int32_t | in | LHS matrix offset(input offset). Range: -127 to 128 |
| active_ch | const int32_t | in | Subset of total_ch processed |
| total_ch | const int32_t | in | Number of channels in LHS/RHS |
| out_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to number of channels. |
| out_mult | const int32_t * | in | Per channel output multiplier. Length of vector is equal to number of channels. |
| out_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| row_x_col | const uint16_t | in | (row_dimension * col_dimension) of LHS/RHS matrix |
| output_bias | const int32_t *const | in | Per channel output bias. Length of vector is equal to number of channels. |
| out | int8_t * | out | Output pointer |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` if an implementation is available or `ARM_CMSIS_NN_NO_IMPL_ERROR` otherwise |

Source: `Include/arm_nnsupportfunctions.h:1492`

## arm_nn_depthwise_conv_s8_planar_candidate

`function` · `c`

```c
static int32_t arm_nn_depthwise_conv_s8_planar_candidate(
    const cmsis_nn_dw_conv_params *dw_conv_params,
    const cmsis_nn_dims *input_dims
)
```

Necessary conditions of the planar rule that are cheap to test inline: at most 32 channels and stride 1. A caller can skip `arm_nn_depthwise_conv_s8_planar()` for layers that fail them without changing which layers it takes.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [1, H, W, C_IN] |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 1 when the layer may take the planar path, 0 when it cannot. |

Source: `Include/arm_nnsupportfunctions.h:1517`

## arm_nn_is_convolve_s8_small_cin

`function` · `c`

```c
static int32_t arm_nn_is_convolve_s8_small_cin(
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    const cmsis_nn_dims *upscale_dims
)
```

The gate of `arm_convolve_s8_small_cin()`: upscale_dims NULL, input depth 1 to 3 with filter depth equal to it, dilation 1, a kernel of at least 1x1 with kernel width x depth at most 16 and at most 48 values, and a positive multiple of 4 output channels. Plain C; it evaluates the same on every build.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
| upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions, or NULL |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 1 when the layer is in the gate, 0 otherwise. |

Source: `Include/arm_nnsupportfunctions.h:1536`

## arm_nn_is_convolve_s8_3x3_c16_s1

`function` · `c`

```c
static int32_t arm_nn_is_convolve_s8_3x3_c16_s1(
    const cmsis_nn_conv_params *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *upscale_dims
)
```

The gate of `arm_convolve_s8_3x3_c16_s1()`: upscale_dims NULL, input and filter depth 16, a 3x3 kernel, and stride and dilation 1. Plain C; it evaluates the same on every build.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params * | in | Convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
| upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions, or NULL |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 1 when the layer is in the gate, 0 otherwise. |

Source: `Include/arm_nnsupportfunctions.h:1562`

## arm_nn_convolve_s8_groups_invalid

`function` · `c`

```c
static int32_t arm_nn_convolve_s8_groups_invalid(
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims
)
```

The group check of `arm_convolve_s8()`, for its direct entries: with groups = C_IN / filter C, C_IN or C_OUT is not a multiple of groups. A filter C of zero or above C_IN gives no group count and is not reported.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 1 when `arm_convolve_s8()` reports the group count as an argument error, 0 otherwise. |

Source: `Include/arm_nnsupportfunctions.h:1582`

## arm_nn_depthwise_conv_s8_planar_bytes

`function` · `c`

```c
int32_t arm_nn_depthwise_conv_s8_planar_bytes(
    const cmsis_nn_dw_conv_params *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims
)
```

Plane size in bytes that `arm_nn_depthwise_conv_s8_planar()` needs for a layer, or -1 when the layer is not one it takes. The rule is plain C and evaluates the same on every build.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [1, H, W, C_IN] |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The plane size in bytes, or -1. |

Source: `Include/arm_nnsupportfunctions.h:1601`

## arm_nn_depthwise_conv_s8_planar

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_depthwise_conv_s8_planar(
    const cmsis_nn_context *ctx,
    const cmsis_nn_context *weight_sum_ctx,
    const cmsis_nn_dw_conv_params *dw_conv_params,
    const cmsis_nn_per_channel_quant_params *quant_params,
    const cmsis_nn_dims *input_dims,
    const int8_t *input,
    const cmsis_nn_dims *filter_dims,
    const int8_t *kernel,
    const cmsis_nn_dims *output_dims,
    int8_t *output
)
```

s8 depthwise convolution with channel multiplier 1 and stride 1, vectorized across the output pixels of one channel plane instead of across channels. It serves the few-channel and 1xk layers of `arm_depthwise_conv_s8_opt()`, with the same scratch buffer and weight sums.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Scratch buffer of `arm_depthwise_conv_s8_opt_get_buffer_size()` bytes |
| weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums from `arm_depthwise_convolve_weight_sum()`, bias included |
| dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
| quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization parameters |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [1, H, W, C_IN] |
| input | const int8_t * | in | Input data pointer |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
| kernel | const int8_t * | in | Filter data pointer |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
| output | int8_t * | out | Output data pointer |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` when the layer was computed, or `ARM_CMSIS_NN_NO_IMPL_ERROR` when it is not one this path takes or its plane does not fit in ctx->size (then nothing is written), or MVE is not available. |

Source: `Include/arm_nnsupportfunctions.h:1626`

## arm_nn_depthwise_conv_nt_t_s4

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_s4(
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t lhs_offset,
    const int32_t active_ch,
    const int32_t total_ch,
    const int32_t *out_shift,
    const int32_t *out_mult,
    const int32_t out_offset,
    const int32_t activation_min,
    const int32_t activation_max,
    const uint16_t row_x_col,
    const int32_t *const output_bias,
    int8_t *out
)
```

Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. rhs consists of packed int4 data. Dimensions are the same for lhs and rhs.

:::note
Tail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `active_ch`.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side matrix |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed). Consists of int4 data packed in an int8 buffer. |
| lhs_offset | const int32_t | in | LHS matrix offset(input offset). Range: -127 to 128 |
| active_ch | const int32_t | in | Subset of total_ch processed |
| total_ch | const int32_t | in | Number of channels in LHS/RHS |
| out_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to number of channels. |
| out_mult | const int32_t * | in | Per channel output multiplier. Length of vector is equal to number of channels. |
| out_offset | const int32_t | in | Offset to be added to the output values. Range: -127 to 128 |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| row_x_col | const uint16_t | in | (row_dimension * col_dimension) of LHS/RHS matrix |
| output_bias | const int32_t *const | in | Per channel output bias. Length of vector is equal to number of channels. |
| out | int8_t * | out | Output pointer |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns one of the two<br><br>- Updated output pointer if an implementation is available - NULL if no implementation is available. |

Source: `Include/arm_nnsupportfunctions.h:1663`

## arm_nn_depthwise_conv_nt_t_s16

`function` · `c`

```c
int16_t * arm_nn_depthwise_conv_nt_t_s16(
    const int16_t *lhs,
    const int8_t *rhs,
    const uint16_t num_ch,
    const int32_t *out_shift,
    const int32_t *out_mult,
    const int32_t activation_min,
    const int32_t activation_max,
    const uint16_t row_x_col,
    const int64_t *const output_bias,
    int16_t *out
)
```

Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. Dimensions are the same for lhs and rhs.

:::note
Tail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `num_ch`.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Input left-hand side matrix |
| rhs | const int8_t * | in | Input right-hand side matrix (transposed) |
| num_ch | const uint16_t | in | Number of channels in LHS/RHS |
| out_shift | const int32_t * | in | Per channel output shift. Length of vector is equal to number of channels. |
| out_mult | const int32_t * | in | Per channel output multiplier. Length of vector is equal to number of channels. |
| activation_min | const int32_t | in | Minimum value to clamp the output to. Range: int8 |
| activation_max | const int32_t | in | Maximum value to clamp the output to. Range: int8 |
| row_x_col | const uint16_t | in | (row_dimension * col_dimension) of LHS/RHS matrix |
| output_bias | const int64_t *const | in | Per channel output bias. Length of vector is equal to number of channels. |
| out | int16_t * | out | Output pointer |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns one of the two<br><br>- Updated output pointer if an implementation is available - NULL if no implementation is available. |

Source: `Include/arm_nnsupportfunctions.h:1699`

## arm_nn_transpose_conv_row_s8_s32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_transpose_conv_row_s8_s32(
    const int8_t *lhs,
    const int8_t *rhs,
    int32_t *output_start,
    const int32_t output_index,
    const int32_t output_max,
    const int32_t rhs_rows,
    const int32_t rhs_cols,
    const int32_t input_channels,
    const int32_t output_channels,
    const int32_t lhs_offset,
    const int32_t row_offset,
    const int32_t input_x,
    const int32_t stride_x,
    const int32_t skip_row_top,
    const int32_t skip_row_bottom
)
```

Row of s8 scalars multiplicated with a s8 matrix ad accumulated into a s32 rolling scratch buffer. Helpfunction for transposed convolution.

:::note
Rolling buffer refers to how the function wraps around the scratch buffer, e.g. it starts writing at [output_start + output_index], writes to [output_start + output_max] and then continues at [output_start] again.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Input left-hand side scalars |
| rhs | const int8_t * | in | Input right-hand side matrix |
| output_start | int32_t * | out | Output buffer start |
| output_index | const int32_t | in | Output buffer current index |
| output_max | const int32_t | in | Output buffer size |
| rhs_rows | const int32_t | in | Number of rows in rhs matrix |
| rhs_cols | const int32_t | in | Number of columns in rhs matrix |
| input_channels | const int32_t | in | Number of input channels |
| output_channels | const int32_t | in | Number of output channels |
| lhs_offset | const int32_t | in | Offset added to lhs before multiplication |
| row_offset | const int32_t | in | Address offset between each row of data output |
| input_x | const int32_t | in | Length of lhs scalar row. |
| stride_x | const int32_t | in | Address offset between each scalar-matrix multiplication result. |
| skip_row_top | const int32_t | in | Skip rows on top of the filter, used for padding. |
| skip_row_bottom | const int32_t | in | Skip rows in the bottom of the filter, used for padding. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns ARM_CMSIS_NN_SUCCESS |

Source: `Include/arm_nnsupportfunctions.h:1735`

## arm_nn_read_q15x2_ia

`function` · `c`

```c
static int32_t arm_nn_read_q15x2_ia(const int16_t **in_q15)
```

Read 2 s16 elements and post increment pointer.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_q15 | const int16_t ** | in, out | Pointer to pointer that holds address of input. Advanced past the elements read. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | q31 value |

Source: `Include/arm_nnsupportfunctions.h:1756`

## arm_nn_read_s8x4_ia

`function` · `c`

```c
static int32_t arm_nn_read_s8x4_ia(const int8_t **in_s8)
```

Read 4 s8 from s8 pointer and post increment pointer.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_s8 | const int8_t ** | in, out | Pointer to pointer that holds address of input. Advanced past the elements read. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | q31 value |

Source: `Include/arm_nnsupportfunctions.h:1771`

## arm_nn_read_s8x2_ia

`function` · `c`

```c
static int32_t arm_nn_read_s8x2_ia(const int8_t **in_s8)
```

Read 2 s8 from s8 pointer and post increment pointer.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_s8 | const int8_t ** | in, out | Pointer to pointer that holds address of input. Advanced past the elements read. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | q31 value |

Source: `Include/arm_nnsupportfunctions.h:1785`

## arm_nn_read_s16x2

`function` · `c`

```c
static int32_t arm_nn_read_s16x2(const int16_t *in)
```

Read 2 int16 values from int16 pointer.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in | const int16_t * | in | pointer to address of input. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | s32 value |

Source: `Include/arm_nnsupportfunctions.h:1799`

## arm_nn_read_s8x4

`function` · `c`

```c
static int32_t arm_nn_read_s8x4(const int8_t *in_s8)
```

Read 4 s8 values.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_s8 | const int8_t * | in | pointer to address of input. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | s32 value |

Source: `Include/arm_nnsupportfunctions.h:1812`

## arm_nn_read_s8x2

`function` · `c`

```c
static int32_t arm_nn_read_s8x2(const int8_t *in_s8)
```

Read 2 s8 values.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_s8 | const int8_t * | in | pointer to address of input. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | s32 value |

Source: `Include/arm_nnsupportfunctions.h:1824`

## arm_nn_write_s8x4_ia

`function` · `c`

```c
static void arm_nn_write_s8x4_ia(int8_t **in, int32_t value)
```

Write four s8 to s8 pointer and increment pointer afterwards.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in | int8_t ** | in, out | Double pointer to destination. Advanced past the bytes written. |
| value | int32_t | in | Four bytes to copy |

Source: `Include/arm_nnsupportfunctions.h:1837`

## arm_memset_s8

`function` · `c`

```c
static void arm_memset_s8(int8_t *dst, const int8_t val, uint32_t block_size)
```

memset optimized for MVE

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int8_t * | in, out | Destination pointer |
| val | const int8_t | in | Value to set |
| block_size | uint32_t | in | Number of bytes to copy. |

Source: `Include/arm_nnsupportfunctions.h:1850`

## arm_memset_s16

`function` · `c`

```c
static void arm_memset_s16(int16_t *dst, const int16_t val, uint32_t block_size)
```

memset optimized for MVE for 16-bit data.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int16_t * | in, out | Destination pointer. |
| val | const int16_t | in | 16-bit value to set. |
| block_size | uint32_t | in | Number of int16_t values to set. |

Source: `Include/arm_nnsupportfunctions.h:1873`

## arm_nn_mat_mult_kernel_s4_s16

`function` · `c`

```c
int8_t * arm_nn_mat_mult_kernel_s4_s16(
    const int8_t *input_a,
    const int16_t *input_b,
    const uint16_t output_ch,
    const int32_t *out_shift,
    const int32_t *out_mult,
    const int32_t out_offset,
    const int32_t activation_min,
    const int32_t activation_max,
    const int32_t num_col_a,
    const int32_t *const output_bias,
    int8_t *out_0
)
```

Matrix-multiplication function for convolution with per-channel requantization and 4 bit weights.

This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_a | const int8_t * | in | pointer to operand A, int8 packed with 2x int4. |
| input_b | const int16_t * | in | pointer to operand B, always consists of 2 vectors. |
| output_ch | const uint16_t | in | number of rows of A |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| activation_min | const int32_t | in | minimum value to clamp the output to. Range : int8 |
| activation_max | const int32_t | in | maximum value to clamp the output to. Range : int8 |
| num_col_a | const int32_t | in | number of columns of A |
| output_bias | const int32_t *const | in | per output channel bias. Range : int32 |
| out_0 | int8_t * | in, out | pointer to output |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns one of the two<br><br>1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |

Source: `Include/arm_nnsupportfunctions.h:2094`

## arm_nn_mat_mult_kernel_s8_s16

`function` · `c`

```c
int8_t * arm_nn_mat_mult_kernel_s8_s16(
    const int8_t *input_a,
    const int16_t *input_b,
    const uint16_t output_ch,
    const int32_t *out_shift,
    const int32_t *out_mult,
    const int32_t out_offset,
    const int16_t activation_min,
    const int16_t activation_max,
    const int32_t num_col_a,
    const int32_t aligned_num_col_a,
    const int32_t *const output_bias,
    int8_t *out_0
)
```

Matrix-multiplication function for convolution with per-channel requantization.

This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_a | const int8_t * | in | pointer to operand A |
| input_b | const int16_t * | in | pointer to operand B, always consists of 2 vectors. |
| output_ch | const uint16_t | in | number of rows of A |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| activation_min | const int16_t | in | minimum value to clamp the output to. Range : int8 |
| activation_max | const int16_t | in | maximum value to clamp the output to. Range : int8 |
| num_col_a | const int32_t | in | number of columns of A |
| aligned_num_col_a | const int32_t | in | number of columns of A aligned by 4 |
| output_bias | const int32_t *const | in | per output channel bias. Range : int32 |
| out_0 | int8_t * | in, out | pointer to output |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns one of the two<br><br>1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |

Source: `Include/arm_nnsupportfunctions.h:2128`

## arm_nn_mat_mult_kernel_row_offset_s8_s16

`function` · `c`

```c
int8_t * arm_nn_mat_mult_kernel_row_offset_s8_s16(
    const int8_t *input_a,
    const int16_t *input_b,
    const uint16_t output_ch,
    const int32_t *out_shift,
    const int32_t *out_mult,
    const int32_t out_offset,
    const int16_t activation_min,
    const int16_t activation_max,
    const int32_t num_col_a,
    const int32_t aligned_num_col_a,
    const int32_t *const output_bias,
    const int32_t row_address_offset,
    int8_t *out_0
)
```

Matrix-multiplication function for convolution with per-channel requantization, supporting an address offset between rows.

This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max.

This function is slighly less performant than arm_nn_mat_mult_kernel_s8_s16, but allows support for grouped convolution. Supported framework: TensorFlow Lite micro.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_a | const int8_t * | in | pointer to operand A |
| input_b | const int16_t * | in | pointer to operand B, always consists of 2 vectors. |
| output_ch | const uint16_t | in | number of rows of A |
| out_shift | const int32_t * | in | pointer to per output channel requantization shift parameter. |
| out_mult | const int32_t * | in | pointer to per output channel requantization multiplier parameter. |
| out_offset | const int32_t | in | output tensor offset. |
| activation_min | const int16_t | in | minimum value to clamp the output to. Range : int8 |
| activation_max | const int16_t | in | maximum value to clamp the output to. Range : int8 |
| num_col_a | const int32_t | in | number of columns of A |
| aligned_num_col_a | const int32_t | in | number of columns of A aligned by 4 |
| output_bias | const int32_t *const | in | per output channel bias. Range : int32 |
| row_address_offset | const int32_t | in | address offset between rows in the output |
| out_0 | int8_t * | in, out | pointer to output |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns one of the two<br><br>1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available. |

Source: `Include/arm_nnsupportfunctions.h:2168`

## arm_nn_softmax_common_s8

`function` · `c`

```c
void arm_nn_softmax_common_s8(
    const int8_t *input,
    const int32_t num_rows,
    const int32_t row_size,
    const int32_t mult,
    const int32_t shift,
    const int32_t diff_min,
    const bool int16_output,
    void *output
)
```

Common softmax function for s8 input and s8 or s16 output.

:::note
Supported framework: TensorFlow Lite micro (bit-accurate)

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const int8_t * | in | Pointer to the input tensor |
| num_rows | const int32_t | in | Number of rows in the input tensor |
| row_size | const int32_t | in | Number of elements in each input row |
| mult | const int32_t | in | Input quantization multiplier |
| shift | const int32_t | in | Input quantization shift within the range [0, 31] |
| diff_min | const int32_t | in | Minimum difference with max in row. Used to check if the quantized exponential operation can be performed |
| int16_output | const bool | in | Indicating s8 output if 0 else s16 output |
| output | void * | out | Pointer to the output tensor |

Source: `Include/arm_nnsupportfunctions.h:2197`

## NN_ROUND

`macro` · `c`

```c
#define NN_ROUND(out_shift) ((0x1 << out_shift) >> 1)
```

macro for adding rounding offset

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| out_shift |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:2210`

## MUL_SAT

`macro` · `c`

```c
#define MUL_SAT(a, b) arm_nn_doubling_high_mult((a), (b))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a |  |  |  |
| b |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:2216`

## MUL_SAT_MVE

`macro` · `c`

```c
#define MUL_SAT_MVE(a, b) arm_doubling_high_mult_mve_32x4((a), (b))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a |  |  |  |
| b |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:2217`

## MUL_POW2

`macro` · `c`

```c
#define MUL_POW2(a, b) arm_nn_mult_by_power_of_two((a), (b))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a |  |  |  |
| b |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:2218`

## DIV_POW2

`macro` · `c`

```c
#define DIV_POW2(a, b) arm_nn_divide_by_power_of_two((a), (b))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a |  |  |  |
| b |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:2220`

## DIV_POW2_MVE

`macro` · `c`

```c
#define DIV_POW2_MVE(a, b) arm_divide_by_power_of_two_mve((a), (b))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a |  |  |  |
| b |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:2221`

## EXP_ON_NEG

`macro` · `c`

```c
#define EXP_ON_NEG(x) arm_nn_exp_on_negative_values((x))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:2223`

## ONE_OVER1

`macro` · `c`

```c
#define ONE_OVER1(x) arm_nn_one_over_one_plus_x_for_x_in_0_1((x))
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:2224`

## arm_nn_doubling_high_mult

`function` · `c`

```c
static int32_t arm_nn_doubling_high_mult(const int32_t m1, const int32_t m2)
```

Saturating doubling high multiply. Result matches NEON instruction VQRDMULH.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| m1 | const int32_t | in | Multiplicand. Range: {NN_Q31_MIN, NN_Q31_MAX} |
| m2 | const int32_t | in | Multiplier. Range: {NN_Q31_MIN, NN_Q31_MAX} |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Result of multiplication. |

Source: `Include/arm_nnsupportfunctions.h:2234`

## arm_nn_doubling_high_mult_no_sat

`function` · `c`

```c
static int32_t arm_nn_doubling_high_mult_no_sat(int32_t m1, int32_t m2)
```

Doubling high multiply without saturation. This is intended for requantization where the scale is a positive integer.

:::note
The result of this matches that of neon instruction VQRDMULH for m1 in range {NN_Q31_MIN, NN_Q31_MAX} and m2 in range {NN_Q31_MIN + 1, NN_Q31_MAX}. Saturation occurs when m1 equals m2 equals NN_Q31_MIN and that is not handled by this function.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| m1 | int32_t | in | Multiplicand. Range: {NN_Q31_MIN, NN_Q31_MAX} |
| m2 | int32_t | in | Multiplier Range: {NN_Q31_MIN, NN_Q31_MAX} |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Result of multiplication. |

Source: `Include/arm_nnsupportfunctions.h:2272`

## arm_nn_divide_by_power_of_two

`function` · `c`

```c
static int32_t arm_nn_divide_by_power_of_two(const int32_t dividend, const int32_t exponent)
```

Rounding divide by power of two.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dividend | const int32_t | in | - Dividend |
| exponent | const int32_t | in | - Divisor = power(2, exponent) Range: [0, 31] |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Rounded result of division. Midpoint is rounded away from zero. |

Source: `Include/arm_nnsupportfunctions.h:2323`

## arm_nn_nonneg_divide_by_pot_s32

`function` · `c`

```c
static int32_t arm_nn_nonneg_divide_by_pot_s32(int32_t dividend, int32_t exponent)
```

Rounding divide by power of two for non-negative values.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dividend | int32_t | in | - Dividend (assumed to be non-negative) |
| exponent | int32_t | in | - Divisor = power(2, exponent) Range: [0, 31] |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Rounded result of division. Midpoint is rounded away from zero. |

Source: `Include/arm_nnsupportfunctions.h:2378`

## arm_nn_requantize

`function` · `c`

```c
static int32_t arm_nn_requantize(const int32_t val, const int32_t multiplier, const int32_t shift)
```

Requantize a given value.

Essentially returns (val * multiplier)/(2 ^ shift) with different rounding depending if CMSIS_NN_USE_SINGLE_ROUNDING is defined or not.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | const int32_t | in | Value to be requantized |
| multiplier | const int32_t | in | Multiplier. Range {NN_Q31_MIN + 1, Q32_MAX} |
| shift | const int32_t | in | Shift. Range: {-31, 30} Default branch: If shift is positive left shift 'val * multiplier' with shift If shift is negative right shift 'val * multiplier' with abs(shift) Single round branch: Input for total_shift in divide by '2 ^ total_shift' |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Default branch: Returns (val * multiplier) with rounding divided by (2 ^ shift) with rounding Single round branch: Returns (val * multiplier)/(2 ^ (31 - shift)) with rounding |

Source: `Include/arm_nnsupportfunctions.h:2416`

## arm_nn_requantize_s64

`function` · `c`

```c
static int32_t arm_nn_requantize_s64(const int64_t val, const int32_t reduced_multiplier, const int32_t shift)
```

Requantize a given 64 bit value.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | const int64_t | in | Value to be requantized in the range {-(1<<47)} to {(1<<47) - 1} |
| reduced_multiplier | const int32_t | in | Reduced multiplier in the range {NN_Q31_MIN + 1, Q32_MAX} to {Q16_MIN + 1, Q16_MAX} |
| shift | const int32_t | in | Left or right shift for 'val * multiplier' in the range {-31} to {7} |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Returns (val * multiplier)/(2 ^ shift) |

Source: `Include/arm_nnsupportfunctions.h:2453`

## arm_nn_sat_lshift_s16

`function` · `c`

```c
static int16_t arm_nn_sat_lshift_s16(int16_t x, int shift)
```

Saturating left shift for int16_t.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | int16_t | in | value to be shifted |
| shift | int | in | Nonpositive values return x; positive values multiply by 2^shift with s16 saturation. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | shifted value |

Source: `Include/arm_nnsupportfunctions.h:2471`

## arm_nn_sqrdmulh_s16

`function` · `c`

```c
static int16_t arm_nn_sqrdmulh_s16(int16_t a, int16_t b)
```

Saturating *Rounding* Doubling High Mul (s16).

Matches NEON SQRDMULH s16

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | int16_t | in | Multiplicand |
| b | int16_t | in | Multiplier |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Result of multiplication. |

Source: `Include/arm_nnsupportfunctions.h:2490`

## arm_nn_sqdmulh_s16

`function` · `c`

```c
static int16_t arm_nn_sqdmulh_s16(int16_t a, int16_t b)
```

Saturating **Non-rounded** Doubling High Mul (s16).

Matches NEON SQDMULH s16

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| a | int16_t | in | Multiplicand |
| b | int16_t | in | Multiplier |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Result of multiplication. |

Source: `Include/arm_nnsupportfunctions.h:2510`

## arm_nn_divide_by_power_of_two_s16

`function` · `c`

```c
static int16_t arm_nn_divide_by_power_of_two_s16(int16_t x, int exponent)
```

Rounding divide by power of two (s16), midpoint away from zero.

Mirrors arm_nn_divide_by_power_of_two() semantics for s16.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | int16_t | in | Dividend |
| exponent | int | in | Divisor = power(2, exponent) Range: [0, 15] |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Rounded result of division. Midpoint is rounded away from zero. |

Source: `Include/arm_nnsupportfunctions.h:2530`

## arm_memcpy_s8

`function` · `c`

```c
static void arm_memcpy_s8(int8_t *dst, const int8_t *src, uint32_t block_size)
```

memcpy optimized for MVE

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int8_t * | in, out | Destination pointer |
| src | const int8_t * | in | Source pointer. |
| block_size | uint32_t | in | Number of bytes to copy. |

Source: `Include/arm_nnsupportfunctions.h:2545`

## arm_memcpy_s16

`function` · `c`

```c
static void arm_memcpy_s16(int16_t *dst, const int16_t *src, uint32_t block_size)
```

memcpy optimized for MVE

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int16_t * | in, out | Destination pointer |
| src | const int16_t * | in | Source pointer. |
| block_size | uint32_t | in | Number of values to copy. |

Source: `Include/arm_nnsupportfunctions.h:2569`

## arm_memcpy_s32

`function` · `c`

```c
static void arm_memcpy_s32(int32_t *dst, const int32_t *src, uint32_t block_size)
```

memcpy optimized for MVE

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int32_t * | in, out | Destination pointer |
| src | const int32_t * | in | Source pointer. |
| block_size | uint32_t | in | Number of values to copy. |

Source: `Include/arm_nnsupportfunctions.h:2581`

## arm_memcpy_q15

`function` · `c`

```c
static void arm_memcpy_q15(int16_t *dst, const int16_t *src, uint32_t block_size)
```

memcpy wrapper for int16

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int16_t * | in, out | Destination pointer |
| src | const int16_t * | in | Source pointer. |
| block_size | uint32_t | in | Number of bytes to copy. |

Source: `Include/arm_nnsupportfunctions.h:2593`

## arm_nn_exp_on_negative_values

`function` · `c`

```c
static int32_t arm_nn_exp_on_negative_values(int32_t val)
```

Fixed-point exp() of a non-positive value.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | int32_t | in | Input in Q5.26 fixed point. Must be less than or equal to 0 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | exp(val) in Q0.31 fixed point. Returns NN_Q31_MAX when `val` is 0. |

Source: `Include/arm_nnsupportfunctions.h:2865`

## SELECT_IF_NON_ZERO

`macro` · `c`

```c
#define SELECT_IF_NON_ZERO(x) { \ mask = MASK_IF_NON_ZERO(remainder & (1 << shift++)); \ result = SELECT_USING_MASK(mask, MUL_SAT(result, x), result); \ }
```

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x |  |  |  |

Source: `Include/arm_nnsupportfunctions.h:2878`

## arm_nn_mult_by_power_of_two

`function` · `c`

```c
static int32_t arm_nn_mult_by_power_of_two(const int32_t val, const int32_t exp)
```

Saturating multiply by a power of two.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | const int32_t | in | Value to be multiplied |
| exp | const int32_t | in | Exponent. Multiplier = power(2, exp) |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | val * 2^exp saturated to the int32 range |

Source: `Include/arm_nnsupportfunctions.h:2905`

## arm_nn_one_over_one_plus_x_for_x_in_0_1

`function` · `c`

```c
static int32_t arm_nn_one_over_one_plus_x_for_x_in_0_1(int32_t val)
```

Fixed-point 1 / (1 + x) for x in [0, 1), computed with Newton-Raphson iterations.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| val | int32_t | in | x in Q0.31 fixed point. Range: [0, NN_Q31_MAX] |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 1 / (1 + x) in Q0.31 fixed point |

Source: `Include/arm_nnsupportfunctions.h:2920`

## arm_nn_write_q15x2_ia

`function` · `c`

```c
static void arm_nn_write_q15x2_ia(int16_t **dest_q15, int32_t src_q31)
```

Write 2 s16 elements and post increment pointer.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dest_q15 | int16_t ** | in, out | Pointer to pointer that holds address of destination. Advanced past the elements written. |
| src_q31 | int32_t | in | Input value to be written. |

Source: `Include/arm_nnsupportfunctions.h:2942`

## arm_nn_write_s8x2_ia

`function` · `c`

```c
static void arm_nn_write_s8x2_ia(int8_t **dst, int16_t src)
```

Write 2 s8 elements and post increment pointer.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | int8_t ** | in, out | Pointer to pointer that holds address of destination. Advanced past the elements written. |
| src | int16_t | in | Input value to be written. |

Source: `Include/arm_nnsupportfunctions.h:2955`

## arm_cmsis_nn_dim_at

`function` · `c`

```c
static int32_t arm_cmsis_nn_dim_at(const cmsis_nn_dims *dims, int32_t index)
```

Get dimension value at specific index.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dims | const cmsis_nn_dims * | in | Pointer to `cmsis_nn_dims` structure |
| index | int32_t | in | Index of dimension to get |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Dimension value at specified index |

Source: `Include/arm_nnsupportfunctions.h:2969`

## arm_cmsis_nn_shape_product

`function` · `c`

```c
static size_t arm_cmsis_nn_shape_product(const int32_t *shape, int32_t length)
```

Calculate the product of all dimensions in a shape array.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| shape | const int32_t * | in | Pointer to array containing shape dimensions |
| length | int32_t | in | Number of dimensions in the shape array |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Product of all dimensions |

Source: `Include/arm_nnsupportfunctions.h:2994`

## arm_nn_lstm_step_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_lstm_step_s8(
    const int8_t *data_in,
    const int8_t *hidden_in,
    int8_t *hidden_out,
    const cmsis_nn_lstm_params *params,
    cmsis_nn_lstm_context *buffers,
    const int32_t batch_offset
)
```

Update LSTM function for an iteration step using s8 input and output, and s16 internally.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const int8_t * | in | Data input pointer |
| hidden_in | const int8_t * | in | Hidden state/ recurrent input pointer |
| hidden_out | int8_t * | out | Hidden state/ recurrent output pointer |
| params | const cmsis_nn_lstm_params * | in | Struct containg all information about the lstm operator, see arm_nn_types. |
| buffers | cmsis_nn_lstm_context * | in, out | Struct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns ARM_CMSIS_NN_SUCCESS |

Source: `Include/arm_nnsupportfunctions.h:3022`

## arm_nn_lstm_step_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_lstm_step_s16(
    const int16_t *data_in,
    const int16_t *hidden_in,
    int16_t *hidden_out,
    const cmsis_nn_lstm_params *params,
    cmsis_nn_lstm_context *buffers,
    const int32_t batch_offset
)
```

Update LSTM function for an iteration step using s16 input and output, and s16 internally.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const int16_t * | in | Data input pointer |
| hidden_in | const int16_t * | in | Hidden state/ recurrent input pointer |
| hidden_out | int16_t * | out | Hidden state/ recurrent output pointer |
| params | const cmsis_nn_lstm_params * | in | Struct containg all information about the lstm operator, see arm_nn_types. |
| buffers | cmsis_nn_lstm_context * | in, out | Struct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns ARM_CMSIS_NN_SUCCESS |

Source: `Include/arm_nnsupportfunctions.h:3046`

## arm_nn_lstm_calculate_gate_s8_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_lstm_calculate_gate_s8_s16(
    const int8_t *data_in,
    const int8_t *hidden_in,
    const cmsis_nn_lstm_gate *gate_data,
    const cmsis_nn_lstm_params *params,
    int16_t *output,
    const int32_t batch_offset
)
```

Updates a LSTM gate for an iteration step of LSTM function, int8x8_16 version.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const int8_t * | in | Data input pointer |
| hidden_in | const int8_t * | in | Hidden state/ recurrent input pointer |
| gate_data | const cmsis_nn_lstm_gate * | in | Struct containing all information about the gate caluclation, see arm_nn_types. |
| params | const cmsis_nn_lstm_params * | in | Struct containing all information about the lstm_operation, see arm_nn_types |
| output | int16_t * | out | Hidden state/ recurrent output pointer |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches, see arm_nn_lstm_step_s8. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns ARM_CMSIS_NN_SUCCESS |

Source: `Include/arm_nnsupportfunctions.h:3067`

## arm_nn_lstm_calculate_gate_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_lstm_calculate_gate_s16(
    const int16_t *data_in,
    const int16_t *hidden_in,
    const cmsis_nn_lstm_gate *gate_data,
    const cmsis_nn_lstm_params *params,
    int16_t *output,
    const int32_t batch_offset
)
```

Updates a LSTM gate for an iteration step of LSTM function, int16x8_16 version.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const int16_t * | in | Data input pointer |
| hidden_in | const int16_t * | in | Hidden state/ recurrent input pointer |
| gate_data | const cmsis_nn_lstm_gate * | in | Struct containing all information about the gate caluclation, see arm_nn_types. |
| params | const cmsis_nn_lstm_params * | in | Struct containing all information about the lstm_operation, see arm_nn_types |
| output | int16_t * | out | Hidden state/ recurrent output pointer |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches, see arm_nn_lstm_step_s16. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns ARM_CMSIS_NN_SUCCESS |

Source: `Include/arm_nnsupportfunctions.h:3088`

## arm_nn_vec_mat_mul_result_acc_s8_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_vec_mat_mul_result_acc_s8_s16(
    const int8_t *lhs,
    const int8_t *rhs,
    const int32_t *effective_bias,
    int16_t *dst,
    const int32_t dst_multiplier,
    const int32_t dst_shift,
    const int32_t rhs_cols,
    const int32_t rhs_rows,
    const int32_t batches,
    const int32_t batch_offset
)
```

The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a "batched" vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int8_t * | in | Batched vector |
| rhs | const int8_t * | in | Weights - input matrix (H(Rows)xW(Columns)) |
| effective_bias | const int32_t * | in | Bias + lhs_offset * kernel_sum term precalculated into a constant vector. |
| dst | int16_t * | out | Output |
| dst_multiplier | const int32_t | in | Multiplier for quantization |
| dst_shift | const int32_t | in | Shift for quantization |
| rhs_cols | const int32_t | in | Vector/matarix column length |
| rhs_rows | const int32_t | in | Row count of matrix |
| batches | const int32_t | in | Batch size |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches in input, see arm_nn_lstm_step_s8. Note that the output is always stored with sequential batches. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:3114`

## arm_nn_vec_mat_mul_result_acc_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_vec_mat_mul_result_acc_s16(
    const int16_t *lhs,
    const int8_t *rhs,
    const int64_t *effective_bias,
    int16_t *dst,
    const int32_t dst_multiplier,
    const int32_t dst_shift,
    const int32_t rhs_cols,
    const int32_t rhs_rows,
    const int32_t batches,
    const int32_t batch_offset
)
```

The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a "batched" vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const int16_t * | in | Batched vector |
| rhs | const int8_t * | in | Weights - input matrix (H(Rows)xW(Columns)) |
| effective_bias | const int64_t * | in | Bias + lhs_offset * kernel_sum term precalculated into a constant vector. |
| dst | int16_t * | out | Output |
| dst_multiplier | const int32_t | in | Multiplier for quantization |
| dst_shift | const int32_t | in | Shift for quantization |
| rhs_cols | const int32_t | in | Vector/matarix column length |
| rhs_rows | const int32_t | in | Row count of matrix |
| batches | const int32_t | in | Batch size |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches in input, see arm_nn_lstm_step_s16. Note that the output is always stored with sequential batches. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns `ARM_CMSIS_NN_SUCCESS` |

Source: `Include/arm_nnsupportfunctions.h:3144`

## arm_elementwise_mul_s16_s8

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_mul_s16_s8(
    const int16_t *input_1_vect,
    const int16_t *input_2_vect,
    int8_t *output,
    const int32_t out_offset,
    const int32_t out_mult,
    const int32_t out_shift,
    const int32_t block_size,
    const int32_t batch_size,
    const int32_t batch_offset
)
```

s16 elementwise multiplication with s8 output

Supported framework: TensorFlow Lite micro

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| output | int8_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| block_size | const int32_t | in | number of samples per batch |
| batch_size | const int32_t | in | number of samples per batch |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches in output, see arm_nn_lstm_step_s8. Note that it is assumed that the input is stored with sequential batches. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns ARM_CMSIS_NN_SUCCESS |

Source: `Include/arm_nnsupportfunctions.h:3171`

## arm_elementwise_mul_s16_batch_offset

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_mul_s16_batch_offset(
    const int16_t *input_1_vect,
    const int16_t *input_2_vect,
    int16_t *output,
    const int32_t out_offset,
    const int32_t out_mult,
    const int32_t out_shift,
    const int32_t block_size,
    const int32_t batch_size,
    const int32_t batch_offset
)
```

s16 elementwise multiplication with s16 output

Supported framework: TensorFlow Lite micro

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| output | int16_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| block_size | const int32_t | in | number of samples per batch |
| batch_size | const int32_t | in | number of samples per batch |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches in output, see arm_nn_lstm_step_s16. Note that it is assumed that the input is stored with sequential batches. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns ARM_CMSIS_NN_SUCCESS |

Source: `Include/arm_nnsupportfunctions.h:3197`

## arm_elementwise_mul_acc_s16

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_mul_acc_s16(
    const int16_t *input_1_vect,
    const int16_t *input_2_vect,
    const int32_t input_1_offset,
    const int32_t input_2_offset,
    int16_t *output,
    const int32_t out_offset,
    const int32_t out_mult,
    const int32_t out_shift,
    const int32_t out_activation_min,
    const int32_t out_activation_max,
    const int32_t block_size
)
```

s16 elementwise multiplication. The result of the multiplication is accumulated to the passed result buffer.

Supported framework: TensorFlow Lite micro

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const int16_t * | in | pointer to input vector 1 |
| input_2_vect | const int16_t * | in | pointer to input vector 2 |
| input_1_offset | const int32_t | in | offset for input 1. Not used. |
| input_2_offset | const int32_t | in | offset for input 2. Not used. |
| output | int16_t * | in, out | pointer to output vector |
| out_offset | const int32_t | in | output offset. Not used. |
| out_mult | const int32_t | in | output multiplier |
| out_shift | const int32_t | in | output shift |
| out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
| out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
| block_size | const int32_t | in | number of samples |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns ARM_CMSIS_NN_SUCCESS |

Source: `Include/arm_nnsupportfunctions.h:3224`

## arm_check_broadcast_required

`function` · `c`

```c
static int32_t arm_check_broadcast_required(const cmsis_nn_dims *shape_1, const cmsis_nn_dims *shape_2)
```

Check if a broadcast is required between 2 `cmsis_nn_dims`.

Compares each dimension and returns 1 if any dimension does not match. This function does not check that broadcast rules are met.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| shape_1 | const cmsis_nn_dims * | in | pointer to input tensor 1 |
| shape_2 | const cmsis_nn_dims * | in | pointer to input tensor 2 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The function returns 1 if a broadcast is required, or 0 if not. |

Source: `Include/arm_nnsupportfunctions.h:3245`

## arm_reduce_get_middle_block_from_arrays

`function` · `c`

```c
static int32_t arm_reduce_get_middle_block_from_arrays(
    const int32_t in_dims,
    const int32_t axis_arr,
    int32_t *outer,
    int32_t *reduce,
    int32_t *inner
)
```

Reports whether the reduced axes of a 4-D tensor form one contiguous block followed by kept axes, as in a NHWC mean over H and W, and gives the flattened sizes. Axes of size 1 are ignored.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in_dims | const int32_t | in | 4-element array {n, h, w, c} |
| axis_arr | const int32_t | in | 4-element mask {axis_n, axis_h, axis_w, axis_c} |
| outer | int32_t * | out | Product of the dims before the reduced block |
| reduce | int32_t * | out | Product of the reduced dims |
| inner | int32_t * | out | Product of the dims after the reduced block |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 1 if the input is [outer, reduce, inner] with the middle dim reduced and inner > 1, otherwise 0 |

Source: `Include/arm_nnsupportfunctions.h:3310`

## ARM_NN_SQRT_S16_TABLEFREE_SHIFT

`macro` · `c`

```c
#define ARM_NN_SQRT_S16_TABLEFREE_SHIFT 14
```

Source: `Include/arm_nnsupportfunctions.h:3364`

## ARM_NN_SQRT_S16_TABLEFREE_MAGIC

`macro` · `c`

```c
#define ARM_NN_SQRT_S16_TABLEFREE_MAGIC UINT32_C(0x5F5FB6C4)
```

Source: `Include/arm_nnsupportfunctions.h:3365`

## ARM_NN_SQRT_S16_TABLEFREE_K0

`macro` · `c`

```c
#define ARM_NN_SQRT_S16_TABLEFREE_K0 (-4.76426697f)
```

Source: `Include/arm_nnsupportfunctions.h:3366`

## ARM_NN_SQRT_S16_TABLEFREE_K1

`macro` · `c`

```c
#define ARM_NN_SQRT_S16_TABLEFREE_K1 (-48.0000114f)
```

Source: `Include/arm_nnsupportfunctions.h:3367`

## arm_nn_sqrt_s16_tablefree_element

`function` · `c`

```c
static int16_t arm_nn_sqrt_s16_tablefree_element(const int32_t value, const float scale)
```

One element of `arm_sqrt_s16_tablefree()`: the float32 chain the MVE path evaluates per lane, so the two agree bit for bit on any IEEE-754 float32 implementation with round-to-nearest-even and a fused multiply-add (fmaf). Every product after the pre-scale either has two uses or feeds an fmaf or a conversion, never another lone multiply, so a compiler allowed to reassociate (-ffast-math) still has no chain to reorder, and no product feeds a bare add, so there is nothing to contract.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| value | const int32_t | in | input code; values <= 0 give 0 |
| scale | const float | in | input_scale / (output_scale * output_scale) as float32 |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | trunc(sqrt(value * scale)) saturated to 32767 |

Source: `Include/arm_nnsupportfunctions.h:3381`
