# heliaCORE.groupSupport

Internal Support functions. Not intended to be called direclty by a CMSIS-NN user.

## arm_nn_exp_poly_coeffs_f32

`attribute` · `c`

```c
const float32_t arm_nn_exp_poly_coeffs_f32[8]
```

Polynomial coefficients used by the float32 MVE exp approximation.

Source: `Include/arm_nnsupportfunctions_flt.h:70`

## arm_nn_exp2_lut_f32

`attribute` · `c`

```c
const float32_t arm_nn_exp2_lut_f32[257]
```

LUT for `2^(i/256)` used by the float32 LUT softmax approximation.

Stores 257 samples for `i = 0..256` so interpolation can safely read `lut[idx + 1]` while indexing the 256 fractional segments.

Source: `Include/arm_nnsupportfunctions_flt.h:78`

## arm_nn_softmax_floor_to_int_f32

`function` · `c`

```c
static int32_t arm_nn_softmax_floor_to_int_f32(float32_t x)
```

Floor of `x` as an int32_t.

Precondition: `x` must already be reduced to the int32_t range and must not be NaN  the float-to-int conversion below is undefined otherwise. The only caller, arm_nn_softmax_exp_lut_f32(), guarantees this by clamping its input to [-80, 80] (NaN included, see there) before scaling by log2(e), which bounds `x` to +/-116.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | in | Value to floor. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Largest int32_t not greater than `x`. |

Source: `Include/arm_nnsupportfunctions_flt.h:92`

## arm_nn_softmax_fp32_from_bits

`function` · `c`

```c
static float32_t arm_nn_softmax_fp32_from_bits(uint32_t bits)
```

Reinterpret a 32-bit pattern as a float32.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bits | uint32_t | in | IEEE-754 binary32 bit pattern. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The float32 value with the bit pattern `bits`. |

Source: `Include/arm_nnsupportfunctions_flt.h:104`

## arm_nn_softmax_exp2i_f32

`function` · `c`

```c
static float32_t arm_nn_softmax_exp2i_f32(int32_t n)
```

Compute `2^n` as a float32 by building the exponent field directly.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| n | int32_t | in | Integer exponent. Clamped to the normal float32 exponent range `[-126, 127]`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `2^n` as a float32. |

Source: `Include/arm_nnsupportfunctions_flt.h:121`

## arm_nn_softmax_exp_taylor_f32

`function` · `c`

```c
static float32_t arm_nn_softmax_exp_taylor_f32(float32_t x)
```

Taylor/Estrin exp approximation for float32 softmax helpers.

The polynomial is evaluated on r in [-ln2/2, ln2/2]. Coefficients come from the Maclaurin series of exp(r): exp(r) ~= 1 + r + r^2/2! + r^3/3! + r^4/4! + r^5/5! + r^6/6! Grouped via Estrin to reduce dependency depth: p = (1 + r) + r^2*(1/2 + r/6) + r^4*(1/24 + r/120) + r^6*(1/720)

Range reduction follows: x = n * ln(2) + r, exp(x) = exp(r) * 2^n

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | in | Exponent argument. Clamped to `[-80, 80]` before evaluation. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Approximation of `exp(x)`. |

Source: `Include/arm_nnsupportfunctions_flt.h:147`

## arm_nn_softmax_exp_lut_f32

`function` · `c`

```c
static float32_t arm_nn_softmax_exp_lut_f32(float32_t x)
```

LUT-based exp approximation for float32 softmax helpers.

Splits `x * log2(e)` into an integer part handled by arm_nn_softmax_exp2i_f32() and a fractional part interpolated linearly from `arm_nn_exp2_lut_f32`.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | in | Exponent argument. Clamped to `[-80, 80]` before evaluation; NaN is flushed to `80`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Approximation of `exp(x)`. |

Source: `Include/arm_nnsupportfunctions_flt.h:184`

## arm_nn_softmax_exp_scalar_f32

`function` · `c`

```c
static float32_t arm_nn_softmax_exp_scalar_f32(float32_t x)
```

Scalar exp approximation used by the float32 softmax paths.

Dispatches to arm_nn_softmax_exp_taylor_f32() when `ARM_NN_USE_EXP_TAYLOR` is defined and to arm_nn_softmax_exp_lut_f32() otherwise.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float32_t | in | Exponent argument. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Approximation of `exp(x)`. |

Source: `Include/arm_nnsupportfunctions_flt.h:254`

## arm_nn_tanh_lut_f32

`attribute` · `c`

```c
const float32_t arm_nn_tanh_lut_f32[385]
```

LUT for tanh(x) sampled over `x in [0, 6]` for float32 helpers.

Stores 385 samples so interpolation can safely read `lut[idx + 1]` while indexing the 384 fractional segments across the interval. The grid spacing (`6/384 == 1/64`) matches the earlier 257-entry `[0, 4]` table, so entries `0..256` are bit-identical to it and the index multiplier is unchanged. Generated by `scripts/gen_tanh_lut_f32.py`.

Source: `Include/arm_nnsupportfunctions_flt.h:276`

## arm_memcpy_f32

`function` · `c`

```c
static void arm_memcpy_f32(float32_t *dst, const float32_t *src, uint32_t block_size)
```

Copy a float32 vector.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | float32_t * | out | Destination buffer. |
| src | const float32_t * | in | Source buffer. |
| block_size | uint32_t | in | Number of elements to copy. |

Source: `Include/arm_nnsupportfunctions_flt.h:311`

## arm_memset_f32

`function` · `c`

```c
static void arm_memset_f32(float32_t *dst, const float32_t val, uint32_t block_size)
```

Set a float32 vector to a constant value.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | float32_t * | out | Destination buffer. |
| val | const float32_t | in | Fill value. |
| block_size | uint32_t | in | Number of elements to write. |

Source: `Include/arm_nnsupportfunctions_flt.h:334`

## arm_nn_depthwise_conv1d_k3_nhwc_f32

`function` · `c`

```c
void arm_nn_depthwise_conv1d_k3_nhwc_f32(
    const float32_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float32_t *kernel,
    const float32_t *b,
    float32_t *out,
    int32_t out_w
)
```

Specialized NHWC depthwise 1D kernel for `k=3`, `ch_mult=1` (float32).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input (and output) channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float32_t * | in | Depthwise weights with shape `[3][in_c]`. |
| b | const float32_t * | in | Optional bias vector of `in_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:367`

## arm_nn_conv1d_k5_nhwc_f32

`function` · `c`

```c
void arm_nn_conv1d_k5_nhwc_f32(
    const float32_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float32_t *kernel,
    const float32_t *b,
    float32_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=5` (float32).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float32_t * | in | Weights with shape `[out_c][5][in_c]`. |
| b | const float32_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |

Source: `Include/arm_nnsupportfunctions_flt.h:387`

## arm_nn_conv1d_k5_packed_f32

`function` · `c`

```c
void arm_nn_conv1d_k5_packed_f32(
    const float32_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float32_t *kernel_packed,
    const float32_t *b,
    float32_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=5` (float32, packed weights).

The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f32`, i.e. `[(5 * in_c)][out_c_block_of_4]`.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float32_t * | in | Weights packed in output-channel blocks of 4 as described above. |
| b | const float32_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |

Source: `Include/arm_nnsupportfunctions_flt.h:411`

## arm_nn_conv1d_k3_nhwc_f32

`function` · `c`

```c
void arm_nn_conv1d_k3_nhwc_f32(
    const float32_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float32_t *kernel,
    const float32_t *b,
    float32_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=3` (float32).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float32_t * | in | Weights with shape `[out_c][3][in_c]`. |
| b | const float32_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:432`

## arm_nn_conv1d_k3_packed_f32

`function` · `c`

```c
void arm_nn_conv1d_k3_packed_f32(
    const float32_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float32_t *kernel_packed,
    const float32_t *b,
    float32_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=3` (float32, packed weights).

The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f32`, i.e. `[(3 * in_c)][out_c_block_of_4]`.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float32_t * | in | Weights packed in output-channel blocks of 4 as described above. |
| b | const float32_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:456`

## arm_nn_maxpool1d_k3s3_nhwc_f32

`function` · `c`

```c
void arm_nn_maxpool1d_k3s3_nhwc_f32(const float32_t *x_nhwc, int32_t in_c, int32_t in_w, float32_t *out, int32_t out_w)
```

Specialized NHWC max-pool 1D kernel for `k=3`, `s=3` (float32).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `3*ow..3*ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:474`

## arm_nn_maxpool1d_k2s2_nhwc_noclip_f32

`function` · `c`

```c
void arm_nn_maxpool1d_k2s2_nhwc_noclip_f32(
    const float32_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    float32_t *out,
    int32_t out_w
)
```

Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` without output clamp (float32).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `2*ow..2*ow+1`. |

Source: `Include/arm_nnsupportfunctions_flt.h:489`

## arm_nn_maxpool1d_k2s2_nhwc_f32

`function` · `c`

```c
void arm_nn_maxpool1d_k2s2_nhwc_f32(
    const float32_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    float32_t *out,
    int32_t out_w,
    float32_t act_min,
    float32_t act_max
)
```

Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` with clamp (float32).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float32_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float32_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `2*ow..2*ow+1`. |
| act_min | float32_t | in | Lower clamp bound applied to `out`. |
| act_max | float32_t | in | Upper clamp bound applied to `out`. |

Source: `Include/arm_nnsupportfunctions_flt.h:506`

## arm_nn_mat_mult_nt_t_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f32(
    const float32_t *lhs,
    const float32_t *rhs,
    const float32_t *bias,
    float32_t *dst,
    int32_t lhs_rows,
    int32_t rhs_rows,
    int32_t rhs_cols,
    int32_t row_address_offset,
    float32_t activation_min,
    float32_t activation_max
)
```

Matrix multiply with non-transposed lhs and transposed rhs rows (float32).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float32_t * | in | Left-hand matrix stored row-major. |
| rhs | const float32_t * | in | Right-hand matrix stored row-major, one row per output channel. |
| bias | const float32_t * | in | Optional bias vector. |
| dst | float32_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of rows in `rhs`. |
| rhs_cols | int32_t | in | Number of columns in `rhs`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float32_t | in | Lower clamp bound. |
| activation_max | float32_t | in | Upper clamp bound. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnsupportfunctions_flt.h:529`

## arm_nn_mat_mult_nt_n_packed_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f32(
    const float32_t *lhs,
    const float32_t *rhs_packed,
    const float32_t *bias,
    float32_t *dst,
    int32_t lhs_rows,
    int32_t rhs_rows,
    int32_t rhs_cols,
    int32_t row_address_offset,
    float32_t activation_min,
    float32_t activation_max
)
```

Matrix multiply with non-transposed lhs and packed non-transposed rhs (float32).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float32_t * | in | Left-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`. |
| rhs_packed | const float32_t * | in | Right-hand matrix with logical shape `[rhs_cols, rhs_rows]`, packed in column blocks of 4. The final block uses the same packed stride and inactive tail lanes are ignored. |
| bias | const float32_t * | in | Optional bias vector. |
| dst | float32_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of logical output columns in the unpacked rhs matrix. |
| rhs_cols | int32_t | in | Shared reduction dimension `K`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float32_t | in | Lower clamp bound. |
| activation_max | float32_t | in | Upper clamp bound. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnsupportfunctions_flt.h:556`

## arm_nn_pack_conv_patch_f32

`function` · `c`

```c
void arm_nn_pack_conv_patch_f32(
    const float32_t *input,
    int32_t in_h,
    int32_t in_w,
    int32_t in_c,
    int32_t kernel_h,
    int32_t kernel_w,
    int32_t stride_h,
    int32_t stride_w,
    int32_t pad_h,
    int32_t pad_w,
    int32_t dilation_h,
    int32_t dilation_w,
    int32_t out_y,
    int32_t out_x,
    float32_t pad_value,
    float32_t *patch_row
)
```

Pack a single convolution patch into one row of a contiguous float32 patch matrix.

Developers familiar with im2row/im2col terminology can think of this as packing one output patch into one row.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Input tensor for one batch in NHWC layout with shape `[in_h][in_w][in_c]`. |
| in_h | int32_t | in | Input height. |
| in_w | int32_t | in | Input width. |
| in_c | int32_t | in | Number of input channels. |
| kernel_h | int32_t | in | Kernel height. |
| kernel_w | int32_t | in | Kernel width. |
| stride_h | int32_t | in | Vertical stride. |
| stride_w | int32_t | in | Horizontal stride. |
| pad_h | int32_t | in | Top padding. |
| pad_w | int32_t | in | Left padding. |
| dilation_h | int32_t | in | Vertical dilation. |
| dilation_w | int32_t | in | Horizontal dilation. |
| out_y | int32_t | in | Output row index of the patch to pack. |
| out_x | int32_t | in | Output column index of the patch to pack. |
| pad_value | float32_t | in | Value written for taps that fall outside the input. |
| patch_row | float32_t * | out | Destination row of `kernel_h * kernel_w * in_c` elements, ordered `[kernel_h][kernel_w][in_c]`. |

Source: `Include/arm_nnsupportfunctions_flt.h:590`

## arm_nn_softmax_1x2_f32

`function` · `c`

```c
void arm_nn_softmax_1x2_f32(const float32_t *in, float32_t *out)
```

Specialized softmax helper for a single float32 row of length 2.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in | const float32_t * | in | Pointer to two contiguous float32 input values. |
| out | float32_t * | out | Pointer to two contiguous float32 output values. |

Source: `Include/arm_nnsupportfunctions_flt.h:613`

## ARM_NN_F16_ACC_BLOCK

`macro` · `c`

```c
#define ARM_NN_F16_ACC_BLOCK (32)
```

Blockwise float16 accumulation on the MVE legs (AmbiqAI/ns-cmsis-nn#586).

A float16 accumulator lane sums at most ARM_NN_F16_ACC_BLOCK taps, in the kernel's tap order, before its partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once. The `_acc16` entries instantiate the same kernel bodies with ARM_NN_F16_ACC_BLOCK_NONE, which never folds.

Source: `Include/arm_nnsupportfunctions_flt.h:626`

## ARM_NN_F16_ACC_BLOCK_NONE

`macro` · `c`

```c
#define ARM_NN_F16_ACC_BLOCK_NONE (INT32_MAX)
```

Source: `Include/arm_nnsupportfunctions_flt.h:627`

## arm_nn_exp_poly_coeffs_f16

`attribute` · `c`

```c
const float32_t arm_nn_exp_poly_coeffs_f16[8]
```

Polynomial coefficients used by the float16 MVE exp approximation.

The float16 MVE helper evaluates the polynomial in widened float32 lanes, but it uses a dedicated coefficient table to keep the float16 path isolated from the float32 feature gate and softmax support stack.

Source: `Include/arm_nnsupportfunctions_flt.h:636`

## arm_nn_exp2_lut_f16

`attribute` · `c`

```c
const uint16_t arm_nn_exp2_lut_f16[257]
```

Quantized binary16 LUT for `2^(i/256)` used by float16 helpers.

Stores 257 samples for `i = 0..256` so interpolation can safely read `lut[idx + 1]` while indexing the 256 fractional segments.

Source: `Include/arm_nnsupportfunctions_flt.h:644`

## arm_nn_tanh_lut_f16

`attribute` · `c`

```c
const uint16_t arm_nn_tanh_lut_f16[257]
```

Quantized binary16 LUT for tanh(x) with `x in [0, 4]`.

Stores 257 samples so interpolation can safely read `lut[idx + 1]` while indexing the 256 fractional segments across the interval.

Source: `Include/arm_nnsupportfunctions_flt.h:652`

## arm_nn_softmax_fp16_from_bits

`function` · `c`

```c
static float16_t arm_nn_softmax_fp16_from_bits(uint16_t bits)
```

Reinterpret a 16-bit pattern as a float16.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bits | uint16_t | in | IEEE-754 binary16 bit pattern. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | The float16 value with the bit pattern `bits`. |

Source: `Include/arm_nnsupportfunctions_flt.h:660`

## arm_nn_softmax_floor_to_int_f16

`function` · `c`

```c
static int32_t arm_nn_softmax_floor_to_int_f16(float16_t x)
```

Floor of `x` as an int32_t.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | in | Value to floor. Must be finite and within the int32_t range. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Largest int32_t not greater than `x`. |

Source: `Include/arm_nnsupportfunctions_flt.h:677`

## arm_nn_softmax_exp2i_f16

`function` · `c`

```c
static float16_t arm_nn_softmax_exp2i_f16(int32_t n)
```

Compute `2^n` as a float16 by building the exponent field directly.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| n | int32_t | in | Integer exponent. Clamped to the normal float16 exponent range `[-14, 15]`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `2^n` as a float16. |

Source: `Include/arm_nnsupportfunctions_flt.h:690`

## arm_nn_softmax_exp_taylor_f16

`function` · `c`

```c
static float16_t arm_nn_softmax_exp_taylor_f16(float16_t x)
```

Taylor/Estrin exp approximation for float16 softmax helpers.

The evaluation uses float32 intermediates to keep the approximation stable, but it is fully independent from the float32 softmax support tables.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | in | Exponent argument. Clamped to `[-80, 80]` before evaluation. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Approximation of `exp(x)`. |

Source: `Include/arm_nnsupportfunctions_flt.h:710`

## arm_nn_softmax_exp_lut_f16

`function` · `c`

```c
static float16_t arm_nn_softmax_exp_lut_f16(float16_t x)
```

LUT-based exp approximation for float16 softmax helpers.

Splits `x * log2(e)` into an integer part handled by arm_nn_softmax_exp2i_f16() and a fractional part interpolated linearly from `arm_nn_exp2_lut_f16`, using float32 intermediates.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | in | Exponent argument. Clamped to `[-80, 80]` before evaluation. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Approximation of `exp(x)`. |

Source: `Include/arm_nnsupportfunctions_flt.h:746`

## arm_nn_softmax_exp_scalar_f16

`function` · `c`

```c
static float16_t arm_nn_softmax_exp_scalar_f16(float16_t x)
```

Scalar exp approximation used by the float16 softmax paths.

Dispatches to arm_nn_softmax_exp_taylor_f16() when `ARM_NN_USE_EXP_TAYLOR` is defined and to arm_nn_softmax_exp_lut_f16() otherwise.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x | float16_t | in | Exponent argument. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Approximation of `exp(x)`. |

Source: `Include/arm_nnsupportfunctions_flt.h:787`

## arm_memcpy_f16

`function` · `c`

```c
static void arm_memcpy_f16(float16_t *dst, const float16_t *src, uint32_t block_size)
```

Copy a float16 vector.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | float16_t * | out | Destination buffer. |
| src | const float16_t * | in | Source buffer. |
| block_size | uint32_t | in | Number of elements to copy. |

Source: `Include/arm_nnsupportfunctions_flt.h:979`

## arm_memset_f16

`function` · `c`

```c
static void arm_memset_f16(float16_t *dst, const float16_t val, uint32_t block_size)
```

Set a float16 vector to a constant value.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dst | float16_t * | out | Destination buffer. |
| val | const float16_t | in | Fill value. |
| block_size | uint32_t | in | Number of elements to write. |

Source: `Include/arm_nnsupportfunctions_flt.h:1002`

## arm_nn_depthwise_conv2x5_nhwc_f16

`function` · `c`

```c
void arm_nn_depthwise_conv2x5_nhwc_f16(
    const float16_t *x_nhwc,
    int32_t batches,
    int32_t in_c,
    int32_t in_w,
    int32_t ch_mult,
    const float16_t *kernel,
    const float16_t *b,
    float16_t *out,
    int32_t out_w,
    float16_t act_min,
    float16_t act_max
)
```

Specialized NHWC depthwise `2x5` kernel (float16).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input tensor in NHWC layout with shape `[batches][2][in_w][in_c]`. |
| batches | int32_t | in | Number of batches. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. |
| ch_mult | int32_t | in | Channel multiplier; the output has `in_c * ch_mult` channels. |
| kernel | const float16_t * | in | Depthwise weights with shape `[2][5][in_c * ch_mult]`. |
| b | const float16_t * | in | Optional bias vector of `in_c * ch_mult` elements. May be NULL. |
| out | float16_t * | out | Output tensor in NHWC layout with shape `[batches][1][out_w][in_c * ch_mult]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input columns `ow..ow+4`. |
| act_min | float16_t | in | Lower clamp bound applied to `out`. |
| act_max | float16_t | in | Upper clamp bound applied to `out`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1039`

## arm_nn_depthwise_conv1d_k3_nhwc_f16

`function` · `c`

```c
void arm_nn_depthwise_conv1d_k3_nhwc_f16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float16_t *kernel,
    const float16_t *b,
    float16_t *out,
    int32_t out_w
)
```

Specialized NHWC depthwise 1D kernel for `k=3`, `ch_mult=1` (float32).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input (and output) channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Depthwise weights with shape `[3][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `in_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1054`

## arm_nn_conv1d_k5_nhwc_f16

`function` · `c`

```c
void arm_nn_conv1d_k5_nhwc_f16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float16_t *kernel,
    const float16_t *b,
    float16_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=5` (float32).

:::note
MVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 5 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 6 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Weights with shape `[out_c][5][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1073`

## arm_nn_conv1d_k5_nhwc_f16_acc16

`function` · `c`

```c
void arm_nn_conv1d_k5_nhwc_f16_acc16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float16_t *kernel,
    const float16_t *b,
    float16_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=5` (float32).

:::note
MVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 5 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 6 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).

:::

:::note
Every MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k5_nhwc_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Weights with shape `[out_c][5][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1088`

## arm_nn_conv1d_k5_packed_f16

`function` · `c`

```c
void arm_nn_conv1d_k5_packed_f16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float16_t *kernel_packed,
    const float16_t *b,
    float16_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=5` (float16, packed weights).

The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(5 * in_c)][out_c_block_of_8]`.

:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 6 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 6) this is the float16-lane result.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float16_t * | in | Weights packed in output-channel blocks of 8 as described above. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1118`

## arm_nn_conv1d_k5_packed_f16_acc16

`function` · `c`

```c
void arm_nn_conv1d_k5_packed_f16_acc16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float16_t *kernel_packed,
    const float16_t *b,
    float16_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=5` (float16, packed weights).

The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(5 * in_c)][out_c_block_of_8]`.

:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 6 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 6) this is the float16-lane result.

:::

:::note
Every MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k5_packed_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float16_t * | in | Weights packed in output-channel blocks of 8 as described above. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+4`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1133`

## arm_nn_conv1d_k3_nhwc_f16

`function` · `c`

```c
void arm_nn_conv1d_k3_nhwc_f16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float16_t *kernel,
    const float16_t *b,
    float16_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=3` (float32).

:::note
MVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 3 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 10 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Weights with shape `[out_c][3][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1153`

## arm_nn_conv1d_k3_nhwc_f16_acc16

`function` · `c`

```c
void arm_nn_conv1d_k3_nhwc_f16_acc16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float16_t *kernel,
    const float16_t *b,
    float16_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=3` (float32).

:::note
MVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 3 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 10 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).

:::

:::note
Every MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k3_nhwc_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel | const float16_t * | in | Weights with shape `[out_c][3][in_c]`. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1168`

## arm_nn_conv1d_k3_packed_f16

`function` · `c`

```c
void arm_nn_conv1d_k3_packed_f16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float16_t *kernel_packed,
    const float16_t *b,
    float16_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=3` (float16, packed weights).

The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(3 * in_c)][out_c_block_of_8]`.

:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 10 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 10) this is the float16-lane result.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float16_t * | in | Weights packed in output-channel blocks of 8 as described above. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1198`

## arm_nn_conv1d_k3_packed_f16_acc16

`function` · `c`

```c
void arm_nn_conv1d_k3_packed_f16_acc16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    const float16_t *kernel_packed,
    const float16_t *b,
    float16_t *out,
    int32_t out_c,
    int32_t out_w
)
```

Specialized NHWC 1D convolution kernel for `k=3` (float16, packed weights).

The packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(3 * in_c)][out_c_block_of_8]`.

:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 10 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 10) this is the float16-lane result.

:::

:::note
Every MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k3_packed_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of input channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| kernel_packed | const float16_t * | in | Weights packed in output-channel blocks of 8 as described above. |
| b | const float16_t * | in | Optional bias vector of `out_c` elements. May be NULL. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][out_c]`. |
| out_c | int32_t | in | Number of output channels. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `ow..ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1213`

## arm_nn_maxpool1d_k3s3_nhwc_f16

`function` · `c`

```c
void arm_nn_maxpool1d_k3s3_nhwc_f16(const float16_t *x_nhwc, int32_t in_c, int32_t in_w, float16_t *out, int32_t out_w)
```

Specialized NHWC max-pool 1D kernel for `k=3`, `s=3` (float16).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `3*ow..3*ow+2`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1227`

## arm_nn_maxpool1d_k2s2_nhwc_noclip_f16

`function` · `c`

```c
void arm_nn_maxpool1d_k2s2_nhwc_noclip_f16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    float16_t *out,
    int32_t out_w
)
```

Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` without output clamp (float16).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `2*ow..2*ow+1`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1238`

## arm_nn_maxpool1d_k2s2_nhwc_f16

`function` · `c`

```c
void arm_nn_maxpool1d_k2s2_nhwc_f16(
    const float16_t *x_nhwc,
    int32_t in_c,
    int32_t in_w,
    float16_t *out,
    int32_t out_w,
    float16_t act_min,
    float16_t act_max
)
```

Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` with clamp (float16).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| x_nhwc | const float16_t * | in | Input row in NHWC layout with shape `[in_w][in_c]`. |
| in_c | int32_t | in | Number of channels. |
| in_w | int32_t | in | Input width. Currently unused by the kernel. |
| out | float16_t * | out | Output row in NHWC layout with shape `[out_w][in_c]`. |
| out_w | int32_t | in | Output width. Output position `ow` reads input positions `2*ow..2*ow+1`. |
| act_min | float16_t | in | Lower clamp bound applied to `out`. |
| act_max | float16_t | in | Upper clamp bound applied to `out`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1249`

## arm_nn_mat_mult_nt_t_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f16(
    const float16_t *lhs,
    const float16_t *rhs,
    const float16_t *bias,
    float16_t *dst,
    int32_t lhs_rows,
    int32_t rhs_rows,
    int32_t rhs_cols,
    int32_t row_address_offset,
    float16_t activation_min,
    float16_t activation_max
)
```

Matrix multiply with non-transposed lhs and transposed rhs rows (float32).

:::note
Accumulation width per leg. MVE legs accumulate blockwise (AmbiqAI/ns-cmsis-nn#586, superseding the float16-lane choice of #417 / #446 for the MVE legs). Up to rhs_cols 32 nothing changes: per-k float16 lanes on the gather path (rhs_cols below the contiguous-K threshold), lane-partial sums then one float16 reduction plus the bias in float16 at rhs_cols 32. Above 32, on the contiguous-K path and the remainder rows, each lane (element k goes to lane k % 8) sums at most 32 of its own taps (256 elements) in float16; each block's lanes are then widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (set by the first block, added to by later ones); the four pair accumulators are summed once as (0+1) + (2+3), so a single block sums ((0+1) + (2+3)) + ((4+5) + (6+7)), the bias is added in float32 and the total rounds to float16 once before the clamp. The float16 reduction up to rhs_cols 32 is ordered by the compiler, which may reorder it under -ffast-math; only the fold's order is fixed. arm_nn_mat_mult_nt_t_f16_acc16 keeps the float16 lanes and float16 reduction throughout. The scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates bias and every product in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457).

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float16_t * | in | Left-hand matrix stored row-major. |
| rhs | const float16_t * | in | Right-hand matrix stored row-major, one row per output channel. |
| bias | const float16_t * | in | Optional bias vector. |
| dst | float16_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of rows in `rhs`. |
| rhs_cols | int32_t | in | Number of columns in `rhs`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float16_t | in | Lower clamp bound. |
| activation_max | float16_t | in | Upper clamp bound. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnsupportfunctions_flt.h:1273`

## arm_nn_mat_mult_nt_t_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f16_acc16(
    const float16_t *lhs,
    const float16_t *rhs,
    const float16_t *bias,
    float16_t *dst,
    int32_t lhs_rows,
    int32_t rhs_rows,
    int32_t rhs_cols,
    int32_t row_address_offset,
    float16_t activation_min,
    float16_t activation_max
)
```

arm_nn_mat_mult_nt_t_f16 with every MVE accumulator lane in float16 (no blockwise fold).

Same arguments, return codes and scalar leg as arm_nn_mat_mult_nt_t_f16; see its accumulation note.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float16_t * | in | Left-hand matrix, row-major `[lhs_rows, rhs_cols]`. |
| rhs | const float16_t * | in | Right-hand matrix, row-major `[rhs_rows, rhs_cols]` (transposed operand). |
| bias | const float16_t * | in | Optional bias vector of `rhs_rows` elements. |
| dst | float16_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of rows in `rhs`. |
| rhs_cols | int32_t | in | Shared reduction dimension `K`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float16_t | in | Lower clamp bound. |
| activation_max | float16_t | in | Upper clamp bound. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnsupportfunctions_flt.h:1301`

## arm_nn_mat_mult_nt_n_packed_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f16(
    const float16_t *lhs,
    const float16_t *rhs_packed,
    const float16_t *bias,
    float16_t *dst,
    int32_t lhs_rows,
    int32_t rhs_rows,
    int32_t rhs_cols,
    int32_t row_address_offset,
    float16_t activation_min,
    float16_t activation_max
)
```

Matrix multiply with non-transposed lhs and packed non-transposed rhs (float16).

:::note
On non-MVE builds the output clamp is the bit-classified scalar clamp of #380, so a NaN accumulator (a NaN in `lhs`, `rhs_packed` or `bias`) propagates to `dst` at every optimization level on the gated toolchains, including the shipped -Ofast. On MVE builds the clamp is vmaxnmq/vminnmq with no NaN restore, so a NaN resolves to a clamp bound there instead.

:::

:::note
Accumulation width per leg: the MVE leg accumulates blockwise (AmbiqAI/ns-cmsis-nn#586): one lane per output column, per-k, the bias opening the first block; every 32 k the float16 partial is widened exactly into per-lane float32 accumulators, which round to float16 once before the clamp (rhs_cols up to 32: exactly the float16-lane result; arm_nn_mat_mult_nt_n_packed_f16_acc16 keeps float16 lanes throughout). The scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates bias and every product in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457).

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float16_t * | in | Left-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`. |
| rhs_packed | const float16_t * | in | Right-hand matrix with logical shape `[rhs_cols, rhs_rows]`, packed in column blocks of 8. The final block uses the same packed stride and inactive tail lanes are ignored. |
| bias | const float16_t * | in | Optional bias vector. |
| dst | float16_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of logical output columns in the unpacked rhs matrix. |
| rhs_cols | int32_t | in | Shared reduction dimension `K`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float16_t | in | Lower clamp bound. |
| activation_max | float16_t | in | Upper clamp bound. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnsupportfunctions_flt.h:1342`

## arm_nn_mat_mult_nt_n_packed_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f16_acc16(
    const float16_t *lhs,
    const float16_t *rhs_packed,
    const float16_t *bias,
    float16_t *dst,
    int32_t lhs_rows,
    int32_t rhs_rows,
    int32_t rhs_cols,
    int32_t row_address_offset,
    float16_t activation_min,
    float16_t activation_max
)
```

arm_nn_mat_mult_nt_n_packed_f16 with every MVE accumulator lane in float16 (no blockwise fold).

Same arguments, return codes and scalar leg as arm_nn_mat_mult_nt_n_packed_f16; see its accumulation note.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lhs | const float16_t * | in | Left-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`. |
| rhs_packed | const float16_t * | in | Right-hand matrix packed in column blocks of 8. |
| bias | const float16_t * | in | Optional bias vector. |
| dst | float16_t * | out | Output matrix. |
| lhs_rows | int32_t | in | Number of rows in `lhs`. |
| rhs_rows | int32_t | in | Number of logical output columns in the unpacked rhs matrix. |
| rhs_cols | int32_t | in | Shared reduction dimension `K`. |
| row_address_offset | int32_t | in | Output row stride, expressed in elements. |
| activation_min | float16_t | in | Lower clamp bound. |
| activation_max | float16_t | in | Upper clamp bound. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnsupportfunctions_flt.h:1370`

## arm_nn_lstm_step_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_lstm_step_f16(
    const float16_t *data_in,
    const float16_t *hidden_in,
    float16_t *hidden_out,
    const cmsis_nn_lstm_params_f16 *params,
    cmsis_nn_lstm_context_f16 *buffers,
    const int32_t batch_offset
)
```

Update LSTM function for an iteration step using float16 input, output and state.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const float16_t * | in | Data input pointer. |
| hidden_in | const float16_t * | in | Hidden state / recurrent input pointer. May be NULL for the first step. |
| hidden_out | float16_t * | out | Hidden state / recurrent output pointer. |
| params | const cmsis_nn_lstm_params_f16 * | in | Struct containing all information about the LSTM operator. |
| buffers | cmsis_nn_lstm_context_f16 * | in, out | Struct containing pointers to mutable cell-state storage. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params/buffers or buffers->cell_state, batch_offset <= 0). |

Source: `Include/arm_nnsupportfunctions_flt.h:1394`

## arm_nn_gru_step_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_gru_step_f16(
    const float16_t *data_in,
    const float16_t *hidden_in,
    float16_t *hidden_out,
    const cmsis_nn_gru_params_f16 *params,
    cmsis_nn_gru_context_f16 *buffers,
    const int32_t batch_offset
)
```

Update GRU function for a single iteration step using float16 data.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const float16_t * | in | Data input pointer for this time step. |
| hidden_in | const float16_t * | in | Recurrent input pointer. NULL for the first step (h_prev = 0). |
| hidden_out | float16_t * | out | Hidden-state output pointer for this time step. |
| params | const cmsis_nn_gru_params_f16 * | in | Struct describing the GRU operator. |
| buffers | cmsis_nn_gru_context_f16 * | in, out | Scratch buffers. temp1 (>= hidden_size) is required when reset_after == 0. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params, batch_offset <= 0, or missing temp1 when reset_after == 0). |

Source: `Include/arm_nnsupportfunctions_flt.h:1414`

## arm_nn_pack_conv_patch_f16

`function` · `c`

```c
void arm_nn_pack_conv_patch_f16(
    const float16_t *input,
    int32_t in_h,
    int32_t in_w,
    int32_t in_c,
    int32_t kernel_h,
    int32_t kernel_w,
    int32_t stride_h,
    int32_t stride_w,
    int32_t pad_h,
    int32_t pad_w,
    int32_t dilation_h,
    int32_t dilation_w,
    int32_t out_y,
    int32_t out_x,
    float16_t pad_value,
    float16_t *patch_row
)
```

Pack a single convolution patch into one row of a contiguous float32 patch matrix.

Developers familiar with im2row/im2col terminology can think of this as packing one output patch into one row.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Input tensor for one batch in NHWC layout with shape `[in_h][in_w][in_c]`. |
| in_h | int32_t | in | Input height. |
| in_w | int32_t | in | Input width. |
| in_c | int32_t | in | Number of input channels. |
| kernel_h | int32_t | in | Kernel height. |
| kernel_w | int32_t | in | Kernel width. |
| stride_h | int32_t | in | Vertical stride. |
| stride_w | int32_t | in | Horizontal stride. |
| pad_h | int32_t | in | Top padding. |
| pad_w | int32_t | in | Left padding. |
| dilation_h | int32_t | in | Vertical dilation. |
| dilation_w | int32_t | in | Horizontal dilation. |
| out_y | int32_t | in | Output row index of the patch to pack. |
| out_x | int32_t | in | Output column index of the patch to pack. |
| pad_value | float16_t | in | Value written for taps that fall outside the input. |
| patch_row | float16_t * | out | Destination row of `kernel_h * kernel_w * in_c` elements, ordered `[kernel_h][kernel_w][in_c]`. |

Source: `Include/arm_nnsupportfunctions_flt.h:1424`

## arm_nn_softmax_1x2_f16

`function` · `c`

```c
void arm_nn_softmax_1x2_f16(const float16_t *in, float16_t *out)
```

Specialized softmax helper for a single float16 row of length 2.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| in | const float16_t * | in | Pointer to two contiguous float16 input values. |
| out | float16_t * | out | Pointer to two contiguous float16 output values. |

Source: `Include/arm_nnsupportfunctions_flt.h:1447`

## arm_nn_lstm_step_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_lstm_step_f32(
    const float32_t *data_in,
    const float32_t *hidden_in,
    float32_t *hidden_out,
    const cmsis_nn_lstm_params_f32 *params,
    cmsis_nn_lstm_context_f32 *buffers,
    const int32_t batch_offset
)
```

Update LSTM function for an iteration step using float32 input, output and state.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const float32_t * | in | Data input pointer. |
| hidden_in | const float32_t * | in | Hidden state / recurrent input pointer. May be NULL for the first step. |
| hidden_out | float32_t * | out | Hidden state / recurrent output pointer. |
| params | const cmsis_nn_lstm_params_f32 * | in | Struct containing all information about the LSTM operator. |
| buffers | cmsis_nn_lstm_context_f32 * | in, out | Struct containing pointers to mutable cell-state storage. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params/buffers or buffers->cell_state, batch_offset <= 0). |

Source: `Include/arm_nnsupportfunctions_flt.h:1466`

## arm_nn_gru_step_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_gru_step_f32(
    const float32_t *data_in,
    const float32_t *hidden_in,
    float32_t *hidden_out,
    const cmsis_nn_gru_params_f32 *params,
    cmsis_nn_gru_context_f32 *buffers,
    const int32_t batch_offset
)
```

Update GRU function for a single iteration step using float32 data.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| data_in | const float32_t * | in | Data input pointer for this time step. |
| hidden_in | const float32_t * | in | Recurrent input pointer. NULL for the first step (h_prev = 0). |
| hidden_out | float32_t * | out | Hidden-state output pointer for this time step. |
| params | const cmsis_nn_gru_params_f32 * | in | Struct describing the GRU operator. |
| buffers | cmsis_nn_gru_context_f32 * | in, out | Scratch buffers. temp1 (>= hidden_size) is required when reset_after == 0. |
| batch_offset | const int32_t | in | Number of timesteps between consecutive batches. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params, batch_offset <= 0, or missing temp1 when reset_after == 0). |

Source: `Include/arm_nnsupportfunctions_flt.h:1486`
