# heliaCORE.FC

Collection of fully-connected and matrix multiplication functions.

Fully-connected layer is basically a matrix-vector multiplication with bias. The matrix is the weights and the input/output vectors are the activation values. Supported {weight, activation} precisions include {8-bit, 8-bit} and {8-bit, 16-bit}

## arm_fully_connected_nhwc_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_fully_connected_nhwc_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_fc_params_f32 *fc_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input,
    const cmsis_nn_dims *filter_dims,
    const float32_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias,
    const cmsis_nn_dims *output_dims,
    float32_t *output
)
```

Fully connected layer, NHWC layout.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f32 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1056`

## arm_fully_connected_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_fully_connected_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_fc_params_f32 *fc_params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input,
    const cmsis_nn_dims *filter_dims,
    const float32_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float32_t *bias,
    const cmsis_nn_dims *output_dims,
    float32_t *output,
    arm_nn_tensor_layout layout
)
```

Fully connected layer, dispatch by layout.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f32 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float32_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float32_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1084`

## arm_fully_connected_f32_get_buffer_size

`function` · `c`

```c
int32_t arm_fully_connected_f32_get_buffer_size(
    const cmsis_nn_fc_params_f32 *fc_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the temporary buffer size required by the fully connected layer.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| fc_params | const cmsis_nn_fc_params_f32 * | in | Fully connected parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:1107`

## arm_batch_matmul_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_batch_matmul_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_bmm_params_f32 *bmm_params,
    const cmsis_nn_dims *input_lhs_dims,
    const float32_t *input_lhs,
    const cmsis_nn_dims *input_rhs_dims,
    const float32_t *input_rhs,
    const cmsis_nn_dims *output_dims,
    float32_t *output
)
```

Batched matrix multiplication.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| bmm_params | const cmsis_nn_bmm_params_f32 * | in | Batch matmul parameters and activation clamp. |
| input_lhs_dims | const cmsis_nn_dims * | in | Left-hand-side input tensor dimensions. |
| input_lhs | const float32_t * | in | Pointer to the left-hand-side input tensor. |
| input_rhs_dims | const cmsis_nn_dims * | in | Right-hand-side input tensor dimensions. |
| input_rhs | const float32_t * | in | Pointer to the right-hand-side input tensor. With `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` each `[K, N]` matrix occupies `K * ceil(N / block) * block` elements (block is 4 for float32, 8 for float16) and consecutive batch matrices are stored back to back at that stride. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1491`

## arm_batch_matmul_f32_get_buffer_size

`function` · `c`

```c
int32_t arm_batch_matmul_f32_get_buffer_size(
    const cmsis_nn_bmm_params_f32 *bmm_params,
    const cmsis_nn_dims *input_lhs_dims,
    const cmsis_nn_dims *input_rhs_dims,
    const cmsis_nn_dims *output_dims
)
```

Get the temporary buffer size required by batched matrix multiplication.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bmm_params | const cmsis_nn_bmm_params_f32 * | in | Batch matmul parameters. |
| input_lhs_dims | const cmsis_nn_dims * | in | Left-hand-side input tensor dimensions. |
| input_rhs_dims | const cmsis_nn_dims * | in | Right-hand-side input tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:1510`

## arm_fully_connected_nhwc_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_fully_connected_nhwc_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_fc_params_f16 *fc_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output
)
```

Fully connected layer, NHWC layout.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3067`

## arm_fully_connected_nhwc_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_fully_connected_nhwc_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_fc_params_f16 *fc_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output
)
```

Fully connected layer, NHWC layout.

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_fully_connected_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_fully_connected_nhwc_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3086`

## arm_fully_connected_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_fully_connected_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_fc_params_f16 *fc_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output,
    arm_nn_tensor_layout layout
)
```

Fully connected layer, dispatch by layout.

:::note
Accumulation width follows the matmul helper the weight format selects (arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16): the scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457); the MVE legs use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3110`

## arm_fully_connected_f16_acc16

`function` · `c`

```c
arm_cmsis_nn_status arm_fully_connected_f16_acc16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_fc_params_f16 *fc_params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *filter_dims,
    const float16_t *kernel,
    const cmsis_nn_dims *bias_dims,
    const float16_t *bias,
    const cmsis_nn_dims *output_dims,
    float16_t *output,
    arm_nn_tensor_layout layout
)
```

Fully connected layer, dispatch by layout.

:::note
Accumulation width follows the matmul helper the weight format selects (arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16): the scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457); the MVE legs use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout.

:::

:::note
Float16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_fully_connected_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_fully_connected_f16.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters and activation clamp. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| kernel | const float16_t * | in | Pointer to the filter tensor data. |
| bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. |
| bias | const float16_t * | in | Optional bias tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor data. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3130`

## arm_fully_connected_f16_get_buffer_size

`function` · `c`

```c
int32_t arm_fully_connected_f16_get_buffer_size(
    const cmsis_nn_fc_params_f16 *fc_params,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *filter_dims,
    const cmsis_nn_dims *output_dims,
    arm_nn_tensor_layout layout
)
```

Get the temporary buffer size required by the fully connected layer.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| fc_params | const cmsis_nn_fc_params_f16 * | in | Fully connected parameters. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:3145`

## arm_batch_matmul_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_batch_matmul_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_bmm_params_f16 *bmm_params,
    const cmsis_nn_dims *input_lhs_dims,
    const float16_t *input_lhs,
    const cmsis_nn_dims *input_rhs_dims,
    const float16_t *input_rhs,
    const cmsis_nn_dims *output_dims,
    float16_t *output
)
```

Batched matrix multiplication.

:::note
Accumulation width. Without adjoints the product goes through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16 and so takes their rule: on the MVE legs a reduction of more than 32 taps per output accumulates blockwise (AmbiqAI/ns-cmsis-nn#586), in float16 up to 32. The adjoint paths accumulate in float16 throughout. There is no `_acc16` entry; a caller that needs float16 lanes on a long reduction calls arm_nn_mat_mult_nt_t_f16_acc16 / arm_nn_mat_mult_nt_n_packed_f16_acc16 per batch.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| bmm_params | const cmsis_nn_bmm_params_f16 * | in | Batch matmul parameters and activation clamp. |
| input_lhs_dims | const cmsis_nn_dims * | in | Left-hand-side input tensor dimensions. |
| input_lhs | const float16_t * | in | Pointer to the left-hand-side input tensor. |
| input_rhs_dims | const cmsis_nn_dims * | in | Right-hand-side input tensor dimensions. |
| input_rhs | const float16_t * | in | Pointer to the right-hand-side input tensor. With `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` each `[K, N]` matrix occupies `K * ceil(N / block) * block` elements (block is 4 for float32, 8 for float16) and consecutive batch matrices are stored back to back at that stride. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3327`

## arm_batch_matmul_f16_get_buffer_size

`function` · `c`

```c
int32_t arm_batch_matmul_f16_get_buffer_size(
    const cmsis_nn_bmm_params_f16 *bmm_params,
    const cmsis_nn_dims *input_lhs_dims,
    const cmsis_nn_dims *input_rhs_dims,
    const cmsis_nn_dims *output_dims
)
```

Get the temporary buffer size required by batched matrix multiplication.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| bmm_params | const cmsis_nn_bmm_params_f16 * | in | Batch matmul parameters. |
| input_lhs_dims | const cmsis_nn_dims * | in | Left-hand-side input tensor dimensions. |
| input_rhs_dims | const cmsis_nn_dims * | in | Right-hand-side input tensor dimensions. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes, or 0 when no scratch buffer is needed. |

Source: `Include/arm_nnfunctions_flt.h:3339`
