# heliaCORE.NNConv

Collection of convolution, depthwise convolution functions and their variants.

The convolution is implemented in 2 steps: im2col and General Matrix Multiplication(GEMM)

im2col is a process of converting each patch of image data into a column. After im2col, the convolution is computed as matrix-matrix multiplication.

To reduce the memory footprint, the im2col is performed partially. Each iteration, only a few column (i.e., patches) are generated followed by GEMM.

## arm_depthwise_nhwc_conv_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_depthwise_nhwc_conv_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input,
    const cmsis_nn_dims *filter_dims,
    const float32_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias,
    const cmsis_nn_dims *output_dims,
    float32_t *output
)
```

Depthwise convolution, NHWC layout.

:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions in NHWC format. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions in NHWC-compatible depthwise format. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions in NHWC format. |
| output | float32_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:79`

## arm_depthwise_conv_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_depthwise_conv_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input,
    const cmsis_nn_dims *filter_dims,
    const float32_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias,
    const cmsis_nn_dims *output_dims,
    float32_t *output,
    arm_nn_tensor_layout layout
)
```

Depthwise convolution, dispatch by layout.

:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:119`

## arm_depthwise_conv_wrapper_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_depthwise_conv_wrapper_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input,
    const cmsis_nn_dims *filter_dims,
    const float32_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias,
    const cmsis_nn_dims *output_dims,
    float32_t *output
)
```

Depthwise convolution wrapper using the CMSIS-NN baseline path.

:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:158`

## arm_depthwise_conv_f32_get_buffer_size

`function` · `c`

```c
int32_t arm_depthwise_conv_f32_get_buffer_size(
    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the temporary buffer size required by depthwise convolution.

:::note
Only one route reads scratch: on MVE builds, an NHWC depthwise with ch_mult != 1, a single input channel and at least CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD output channels can run as a regular convolution, which needs the repacked filter, `ROUND_UP(output_dims->c, 4) * filter_dims->h * filter_dims->w * sizeof(float32_t)` bytes (`ROUND_UP(output_dims->c, 8)` and `sizeof(float16_t)` for `_f16`), plus `arm_convolve_wrapper_f32_get_buffer_size` (`_f16`) for that convolution. The query reserves that for every such layer; one an exact-shape specialization takes first runs without it, so the size is an upper bound there. Every other layer  the `ch_mult == 1` direct kernel and the generic kernel  runs without scratch and the query returns 0 (AmbiqAI/ns-cmsis-nn#448, #625).

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:189`

## arm_depthwise_conv_wrapper_f32_get_buffer_size

`function` · `c`

```c
int32_t arm_depthwise_conv_wrapper_f32_get_buffer_size(
    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims
)
```

Get the buffer size required by the depthwise convolution wrapper.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params_f32 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:205`

## arm_convolve_nhwc_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_nhwc_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data
)
```

Convolution, NHWC layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float32_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:232`

## arm_convolve_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data,
    arm_nn_tensor_layout layout
)
```

Convolution, dispatch by layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:266`

## arm_convolve_wrapper_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_wrapper_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data
)
```

Convolution wrapper using the CMSIS-NN baseline path.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float32_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:294`

## arm_convolve_1x1_nhwc_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1x1_nhwc_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data
)
```

1x1 convolution, NHWC layout.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float32_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:321`

## arm_convolve_1x1_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1x1_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data,
    arm_nn_tensor_layout layout
)
```

1x1 convolution, dispatch by layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:354`

## arm_convolve_1_x_n_nhwc_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data
)
```

1xN convolution, NHWC layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float32_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:391`

## arm_convolve_1_x_n_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1_x_n_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data,
    arm_nn_tensor_layout layout
)
```

1xN convolution, dispatch by layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:428`

## arm_convolve_f32_get_buffer_size

`function` · `c`

```c
int32_t arm_convolve_f32_get_buffer_size(
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the temporary buffer size required by convolution.

:::note
When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still reports only the temporary input/im2col scratch requirement. Any offline-packed filter storage is expected to be provided by the caller.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:456`

## arm_convolve_wrapper_f32_get_buffer_size

`function` · `c`

```c
int32_t arm_convolve_wrapper_f32_get_buffer_size(
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims
)
```

Get the buffer size required by the convolution wrapper.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:472`

## arm_convolve_1x1_f32_get_buffer_size

`function` · `c`

```c
int32_t arm_convolve_1x1_f32_get_buffer_size(
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the buffer size required by 1x1 convolution.

:::note
Returns `0` for the unity-stride no-pack path. For non-unity-stride NHWC 1x1 convolution, the returned scratch size enables the packed-tile + GEMM path. When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still excludes the offline-packed filter storage itself.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:493`

## arm_convolve_1_x_n_f32_get_buffer_size

`function` · `c`

```c
int32_t arm_convolve_1_x_n_f32_get_buffer_size(
    const cmsis_nn_conv_params_f32 *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the buffer size required by 1xN convolution.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f32 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:510`

## arm_transpose_conv_wrapper_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_transpose_conv_wrapper_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_context *output_ctx,
    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data,
    arm_nn_tensor_layout layout
)
```

Transpose convolution wrapper using the CMSIS-NN baseline path.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1540`

## arm_transpose_conv_nhwc_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_transpose_conv_nhwc_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_context *output_ctx,
    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data
)
```

Transpose convolution, NHWC layout.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float32_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1570`

## arm_transpose_conv_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_transpose_conv_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_context *output_ctx,
    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float32_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float32_t *output_data,
    arm_nn_tensor_layout layout
)
```

Transpose convolution, dispatch by layout.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1600`

## arm_transpose_conv_f32_get_buffer_size

`function` · `c`

```c
int32_t arm_transpose_conv_f32_get_buffer_size(
    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *out_dims
)
```

Get the temporary buffer size required by transpose convolution.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| out_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:1623`

## arm_transpose_conv_f32_get_reverse_conv_buffer_size

`function` · `c`

```c
int32_t arm_transpose_conv_f32_get_reverse_conv_buffer_size(
    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims
)
```

Get the reverse-convolution workspace size used by transpose convolution helpers.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f32 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no reverse-convolution buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:1638`

## arm_depthwise_nhwc_conv_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_depthwise_nhwc_conv_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output
)
```

Depthwise convolution, NHWC layout.

:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions in NHWC format. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions in NHWC-compatible depthwise format. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions in NHWC format. |
| output | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2224`

## arm_depthwise_nhwc_conv_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_depthwise_nhwc_conv_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output
)
```

Depthwise convolution, NHWC layout.

:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.

:::

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_nhwc_conv_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_nhwc_conv_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions in NHWC format. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions in NHWC-compatible depthwise format. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions in NHWC format. |
| output | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2243`

## arm_depthwise_conv_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_depthwise_conv_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output,
    arm_nn_tensor_layout layout
)
```

Depthwise convolution, dispatch by layout.

:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2273`

## arm_depthwise_conv_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_depthwise_conv_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output,
    arm_nn_tensor_layout layout
)
```

Depthwise convolution, dispatch by layout.

:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.

:::

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_conv_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_conv_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2293`

## arm_depthwise_conv_wrapper_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_depthwise_conv_wrapper_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output
)
```

Depthwise convolution wrapper using the CMSIS-NN baseline path.

:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2324`

## arm_depthwise_conv_wrapper_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_depthwise_conv_wrapper_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output
)
```

Depthwise convolution wrapper using the CMSIS-NN baseline path.

:::note
When `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.

:::

:::note
Accumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.

:::

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_conv_wrapper_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_conv_wrapper_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2343`

## arm_depthwise_conv_f16_get_buffer_size

`function` · `c`

```c
int32_t arm_depthwise_conv_f16_get_buffer_size(
    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the temporary buffer size required by depthwise convolution.

:::note
Only one route reads scratch: on MVE builds, an NHWC depthwise with ch_mult != 1, a single input channel and at least CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD output channels can run as a regular convolution, which needs the repacked filter, `ROUND_UP(output_dims->c, 4) * filter_dims->h * filter_dims->w * sizeof(float32_t)` bytes (`ROUND_UP(output_dims->c, 8)` and `sizeof(float16_t)` for `_f16`), plus `arm_convolve_wrapper_f32_get_buffer_size` (`_f16`) for that convolution. The query reserves that for every such layer; one an exact-shape specialization takes first runs without it, so the size is an upper bound there. Every other layer  the `ch_mult == 1` direct kernel and the generic kernel  runs without scratch and the query returns 0 (AmbiqAI/ns-cmsis-nn#448, #625).

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:2357`

## arm_depthwise_conv_wrapper_f16_get_buffer_size

`function` · `c`

```c
int32_t arm_depthwise_conv_wrapper_f16_get_buffer_size(
    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims
)
```

Get the buffer size required by the depthwise convolution wrapper.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| dw_conv_params | const cmsis_nn_dw_conv_params_f16 * | in | Depthwise convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:2366`

## arm_convolve_nhwc_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_nhwc_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data
)
```

Convolution, NHWC layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2374`

## arm_convolve_nhwc_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_nhwc_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data
)
```

Convolution, NHWC layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.

:::

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_nhwc_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2393`

## arm_convolve_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data,
    arm_nn_tensor_layout layout
)
```

Convolution, dispatch by layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.

:::

:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct OHWI / NT_N_PACKED fallback accumulates bias and every tap in float32 and rounds to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #457); the 1x1, 1xN and patch-GEMM paths go through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16, whose scalar legs do the same, as do the 1xN no-padding OHWI region and the k=3 / k=5 conv1d specializations (#465). MVE leg: the direct small-C kernel accumulates in float32 (widened lanes); the direct OHWI / NT_N_PACKED fallback, every matmul-backed path (1x1, 1xN, patch-GEMM), the 1xN no-padding region and the conv1d specializations use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. Where a dot product spreads its taps over the lanes of one vector (OHWI rows, the contiguous-K matmul, the conv1d k=3 / k=5 OHWI kernels), each lane's own taps form its blocks and, once the reduction exceeds 32 taps, each block's lanes are widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (the first block sets it); the four pair accumulators are summed once as (0+1) + (2+3), the bias is added in float32 and the total rounds once. The k=3 / k=5 kernels close a block on a whole input-channel step (30 taps per lane). An output's taps are the ones its kernel multiplies: the direct fallback skips padded taps, so an edge output counts only its in-range taps, while patch-GEMM and the 1xN padded regions multiply a zero-padded patch and count its padded taps too. Patch-GEMM runs only when ctx provides its scratch, so an edge output's value can depend on whether ctx->buf is given. The fold's order is fixed; the float16 reduction of a dot of at most 32 taps is left to the compiler, which may reorder it under -ffast-math, as before #586.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2431`

## arm_convolve_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data,
    arm_nn_tensor_layout layout
)
```

Convolution, dispatch by layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.

:::

:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct OHWI / NT_N_PACKED fallback accumulates bias and every tap in float32 and rounds to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #457); the 1x1, 1xN and patch-GEMM paths go through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16, whose scalar legs do the same, as do the 1xN no-padding OHWI region and the k=3 / k=5 conv1d specializations (#465). MVE leg: the direct small-C kernel accumulates in float32 (widened lanes); the direct OHWI / NT_N_PACKED fallback, every matmul-backed path (1x1, 1xN, patch-GEMM), the 1xN no-padding region and the conv1d specializations use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. Where a dot product spreads its taps over the lanes of one vector (OHWI rows, the contiguous-K matmul, the conv1d k=3 / k=5 OHWI kernels), each lane's own taps form its blocks and, once the reduction exceeds 32 taps, each block's lanes are widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (the first block sets it); the four pair accumulators are summed once as (0+1) + (2+3), the bias is added in float32 and the total rounds once. The k=3 / k=5 kernels close a block on a whole input-channel step (30 taps per lane). An output's taps are the ones its kernel multiplies: the direct fallback skips padded taps, so an edge output counts only its in-range taps, while patch-GEMM and the 1xN padded regions multiply a zero-padded patch and count its padded taps too. Patch-GEMM runs only when ctx provides its scratch, so an edge output's value can depend on whether ctx->buf is given. The fold's order is fixed; the float16 reduction of a dot of at most 32 taps is left to the compiler, which may reorder it under -ffast-math, as before #586.

:::

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2451`

## arm_convolve_wrapper_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_wrapper_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data
)
```

Convolution wrapper using the CMSIS-NN baseline path.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2466`

## arm_convolve_wrapper_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_wrapper_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data
)
```

Convolution wrapper using the CMSIS-NN baseline path.

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_wrapper_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_wrapper_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2485`

## arm_convolve_1x1_nhwc_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1x1_nhwc_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data
)
```

1x1 convolution, NHWC layout.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2499`

## arm_convolve_1x1_nhwc_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1x1_nhwc_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data
)
```

1x1 convolution, NHWC layout.

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1x1_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1x1_nhwc_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2518`

## arm_convolve_1x1_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1x1_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data,
    arm_nn_tensor_layout layout
)
```

1x1 convolution, dispatch by layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2532`

## arm_convolve_1x1_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1x1_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data,
    arm_nn_tensor_layout layout
)
```

1x1 convolution, dispatch by layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.

:::

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1x1_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1x1_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2552`

## arm_convolve_1_x_n_nhwc_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data
)
```

1xN convolution, NHWC layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2567`

## arm_convolve_1_x_n_nhwc_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data
)
```

1xN convolution, NHWC layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.

:::

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1_x_n_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1_x_n_nhwc_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN]. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN]. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT]. |
| output_data | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2586`

## arm_convolve_1_x_n_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1_x_n_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data,
    arm_nn_tensor_layout layout
)
```

1xN convolution, dispatch by layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.

:::

:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and every product accumulate in float32 and round to float16 once (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: the padded regions go through the matmul helpers and the no-padding region through a strided kernel, all with blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. The padded regions multiply a zero-padded patch row, so their outputs count the padded taps as well.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2610`

## arm_convolve_1_x_n_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_convolve_1_x_n_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data,
    arm_nn_tensor_layout layout
)
```

1xN convolution, dispatch by layout.

:::note
When `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.

:::

:::note
Accumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and every product accumulate in float32 and round to float16 once (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: the padded regions go through the matmul helpers and the no-padding region through a strided kernel, all with blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. The padded regions multiply a zero-padded patch row, so their outputs count the padded taps as well.

:::

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1_x_n_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1_x_n_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters (stride, padding, dilation and activation clamp). |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format depends on `layout`. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format depends on `layout`. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT]. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format depends on `layout`. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2630`

## arm_convolve_f16_get_buffer_size

`function` · `c`

```c
int32_t arm_convolve_f16_get_buffer_size(
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the temporary buffer size required by convolution.

:::note
When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still reports only the temporary input/im2col scratch requirement. Any offline-packed filter storage is expected to be provided by the caller.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:2645`

## arm_convolve_wrapper_f16_get_buffer_size

`function` · `c`

```c
int32_t arm_convolve_wrapper_f16_get_buffer_size(
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims
)
```

Get the buffer size required by the convolution wrapper.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:2654`

## arm_convolve_1x1_f16_get_buffer_size

`function` · `c`

```c
int32_t arm_convolve_1x1_f16_get_buffer_size(
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the buffer size required by 1x1 convolution.

:::note
Returns `0` for the unity-stride no-pack path. For non-unity-stride NHWC 1x1 convolution, the returned scratch size enables the packed-tile + GEMM path. When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still excludes the offline-packed filter storage itself.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:2662`

## arm_convolve_1_x_n_f16_get_buffer_size

`function` · `c`

```c
int32_t arm_convolve_1_x_n_f16_get_buffer_size(
    const cmsis_nn_conv_params_f16 *conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the buffer size required by 1xN convolution.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| conv_params | const cmsis_nn_conv_params_f16 * | in | Convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:2671`

## arm_transpose_conv_wrapper_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_transpose_conv_wrapper_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_context *output_ctx,
    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data,
    arm_nn_tensor_layout layout
)
```

Transpose convolution wrapper using the CMSIS-NN baseline path.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3354`

## arm_transpose_conv_nhwc_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_transpose_conv_nhwc_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_context *output_ctx,
    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data
)
```

Transpose convolution, NHWC layout.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3370`

## arm_transpose_conv_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_transpose_conv_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_context *output_ctx,
    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input_data,
    const cmsis_nn_dims *filter_dims,
    const float16_t *filter_data,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias_data,
    const cmsis_nn_dims *output_dims,
    float16_t *output_data,
    arm_nn_tensor_layout layout
)
```

Transpose convolution, dispatch by layout.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| output_ctx | const cmsis_nn_context * | in | Output context. Unused; may be NULL. |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input_data | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| filter_data | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias_data | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output_data | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3385`

## arm_transpose_conv_f16_get_buffer_size

`function` · `c`

```c
int32_t arm_transpose_conv_f16_get_buffer_size(
    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *out_dims
)
```

Get the temporary buffer size required by transpose convolution.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| out_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:3401`

## arm_transpose_conv_f16_get_reverse_conv_buffer_size

`function` · `c`

```c
int32_t arm_transpose_conv_f16_get_reverse_conv_buffer_size(
    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims
)
```

Get the reverse-convolution workspace size used by transpose convolution helpers.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| transpose_conv_params | const cmsis_nn_transpose_conv_params_f16 * | in | Transpose convolution parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no reverse-convolution buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:3410`
