# heliaCORE.NNSupport

## arm_transpose_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_transpose_f32(
    const cmsis_nn_context *ctx,
    const cmsis_nn_transpose_params_f32 *params,
    const cmsis_nn_dims *input_dims,
    const float32_t *input,
    const cmsis_nn_dims *output_dims,
    float32_t *output
)
```

Transpose a floating-point tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| params | const cmsis_nn_transpose_params_f32 * | in | Transpose parameters, including permutation and layout information. num_dims must be in [1, 4] and perm must be a bijection over [0, num_dims - 1]. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float32_t * | in | Pointer to the input tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. The first params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[perm[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not. |
| output | float32_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1135`

## arm_concatenation_f32_x

`function` · `c`

```c
void arm_concatenation_f32_x(
    const float32_t *input,
    int32_t input_x,
    int32_t input_y,
    int32_t input_z,
    int32_t input_w,
    float32_t *output,
    int32_t output_x,
    uint32_t offset_x
)
```

Concatenate tensors along the X axis.

Call once per input tensor: `offset_x` selects where the input is stored along the X axis of the output tensor and must be advanced by `input_x` after each call. The output tensor must have the same height, channels and batch size as every input tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float32_t * | out | Pointer to the output tensor. |
| output_x | int32_t | in | Width of the output tensor. |
| offset_x | uint32_t | in | Offset on the X axis at which the input tensor is stored. Must be less than `output_x`. |

Source: `Include/arm_nnfunctions_flt.h:1158`

## arm_concatenation_f32_y

`function` · `c`

```c
void arm_concatenation_f32_y(
    const float32_t *input,
    int32_t input_x,
    int32_t input_y,
    int32_t input_z,
    int32_t input_w,
    float32_t *output,
    int32_t output_y,
    uint32_t offset_y
)
```

Concatenate tensors along the Y axis.

Call once per input tensor: `offset_y` selects where the input is stored along the Y axis of the output tensor and must be advanced by `input_y` after each call. The output tensor must have the same width, channels and batch size as every input tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float32_t * | out | Pointer to the output tensor. |
| output_y | int32_t | in | Height of the output tensor. |
| offset_y | uint32_t | in | Offset on the Y axis at which the input tensor is stored. Must be less than `output_y`. |

Source: `Include/arm_nnfunctions_flt.h:1183`

## arm_concatenation_f32_z

`function` · `c`

```c
void arm_concatenation_f32_z(
    const float32_t *input,
    int32_t input_x,
    int32_t input_y,
    int32_t input_z,
    int32_t input_w,
    float32_t *output,
    int32_t output_z,
    uint32_t offset_z
)
```

Concatenate tensors along the Z axis.

Call once per input tensor: `offset_z` selects where the input is stored along the Z axis of the output tensor and must be advanced by `input_z` after each call. The output tensor must have the same width, height and batch size as every input tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float32_t * | out | Pointer to the output tensor. |
| output_z | int32_t | in | Channels in the output tensor. |
| offset_z | uint32_t | in | Offset on the Z axis at which the input tensor is stored. Must be less than `output_z`. |

Source: `Include/arm_nnfunctions_flt.h:1208`

## arm_concatenation_f32_w

`function` · `c`

```c
void arm_concatenation_f32_w(
    const float32_t *input,
    int32_t input_x,
    int32_t input_y,
    int32_t input_z,
    int32_t input_w,
    float32_t *output,
    uint32_t offset_w
)
```

Concatenate tensors along the W axis.

Call once per input tensor: `offset_w` selects where the input is stored along the W axis of the output tensor and must be advanced by `input_w` after each call. The output tensor must have the same width, height and channels as every input tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float32_t * | out | Pointer to the output tensor. |
| offset_w | uint32_t | in | Offset on the W axis at which the input tensor is stored. |

Source: `Include/arm_nnfunctions_flt.h:1232`

## arm_concatenation_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_concatenation_f32(
    const float32_t *const *input_data,
    int32_t num_inputs,
    const int32_t *axis_sizes,
    int32_t output_dims,
    const int32_t *output_shape,
    int32_t axis,
    float32_t *output_data
)
```

Concatenate float32 tensors of any rank along one axis.

Rank-agnostic sibling of the 4-D per-axis arm_concatenation_f32_{x,y,z,w} entry points: all inputs at once, any rank, any axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Input `s` has the output shape with `output_shape`[axis] replaced by `axis_sizes`[s]; the inputs are laid down in order along the axis. Inputs must not overlap the output. A dimension of 0 is accepted and copies nothing.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t *const * | in | Array of `num_inputs` pointers to the flattened (row-major) inputs. |
| num_inputs | int32_t | in | Number of inputs (>= 1). |
| axis_sizes | const int32_t * | in | Array of length `num_inputs:` each input's extent along `axis`. |
| output_dims | int32_t | in | Number of dimensions in `output_shape` (>= 1). |
| output_shape | const int32_t * | in | Output shape; `output_shape`[axis] must equal the sum of `axis_sizes`. |
| axis | int32_t | in | Axis to concatenate along (0 <= axis < output_dims). |
| output_data | float32_t * | out | Pointer to the flattened output. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, size entry, size sum, NULL pointer or an element count above INT32_MAX. |

Source: `Include/arm_nnfunctions_flt.h:1259`

## arm_split_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_split_f32(
    const float32_t *input_data,
    int32_t input_dims,
    const int32_t *input_shape,
    int32_t axis,
    int32_t num_splits,
    const int32_t *split_dims,
    float32_t *const *output_data
)
```

Split a float32 tensor of any rank into several tensors along one axis.

Inverse of arm_concatenation_f32; per-split lengths also cover SPLIT_V. Output `s` has the input shape with `input_shape`[axis] replaced by `split_dims`[s]. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input. A dimension of 0 is accepted and copies nothing.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Pointer to the flattened (row-major) input. |
| input_dims | int32_t | in | Number of dimensions in `input_shape` (>= 1). |
| input_shape | const int32_t * | in | Input shape; `input_shape`[axis] must equal the sum of `split_dims`. |
| axis | int32_t | in | Axis to split along (0 <= axis < input_dims). |
| num_splits | int32_t | in | Number of outputs (>= 1). |
| split_dims | const int32_t * | in | Array of length `num_splits:` each output's extent along `axis`. |
| output_data | float32_t *const * | out | Array of `num_splits` pointers to the flattened outputs. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, split entry, split sum, NULL pointer or an element count above INT32_MAX. |

Source: `Include/arm_nnfunctions_flt.h:1285`

## arm_pack_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_pack_f32(
    const float32_t *const *input_data,
    int32_t num_inputs,
    int32_t input_dims,
    const int32_t *input_shape,
    int32_t axis,
    float32_t *output_data
)
```

Stack float32 tensors of equal shape along a new axis (TFLite PACK).

The output shape is `input_shape` with `num_inputs` inserted at `axis`; input `s` lands at index `s` of that axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Inputs must not overlap the output. Rank-0 inputs (`input_dims` == 0, `axis` == 0) stack into a vector.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t *const * | in | Array of `num_inputs` pointers to the flattened (row-major) inputs. |
| num_inputs | int32_t | in | Number of inputs (>= 1). |
| input_dims | int32_t | in | Number of dimensions of each input (>= 0). |
| input_shape | const int32_t * | in | Shape shared by every input (may be NULL when `input_dims` is 0). |
| axis | int32_t | in | Position of the new axis in the output (0 <= axis <= input_dims). |
| output_data | float32_t * | out | Pointer to the flattened output. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, NULL pointer or an element count above INT32_MAX. |

Source: `Include/arm_nnfunctions_flt.h:1310`

## arm_unpack_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_unpack_f32(
    const float32_t *input_data,
    int32_t input_dims,
    const int32_t *input_shape,
    int32_t axis,
    float32_t *const *output_data
)
```

Unstack a float32 tensor along one axis into `input_shape`[axis] tensors (TFLite UNPACK).

Inverse of arm_pack_f32: output `s` is the input with the axis fixed at index `s` and removed from the shape. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Pointer to the flattened (row-major) input. |
| input_dims | int32_t | in | Number of dimensions in `input_shape` (>= 1). |
| input_shape | const int32_t * | in | Input shape; `input_shape`[axis] (>= 1) is the number of outputs. |
| axis | int32_t | in | Axis to unstack (0 <= axis < input_dims). |
| output_data | float32_t *const * | out | Array of `input_shape`[axis] pointers to the flattened outputs. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, a zero-extent unstack axis (no outputs to produce), NULL pointer or an element count above INT32_MAX. |

Source: `Include/arm_nnfunctions_flt.h:1334`

## arm_batch_norm_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_batch_norm_f32(
    const float32_t *input,
    float32_t *output,
    const float32_t *scale,
    const float32_t *bias,
    const cmsis_nn_dims *input_dims,
    arm_nn_tensor_layout layout
)
```

Apply batch normalization.

Computes `output = input * scale[c] + bias[c]` for every element of channel `c`, with `scale` and `bias` holding the pre-folded per-channel factors.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor data. Format: [N, H, W, C]. |
| output | float32_t * | out | Pointer to the output tensor data, same shape as `input`. |
| scale | const float32_t * | in | Per-channel scale, `input_dims->c` values. |
| bias | const float32_t * | in | Per-channel bias, `input_dims->c` values. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Every dimension must be positive. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Must be `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1390`

## arm_reshape_f32

`function` · `c`

```c
void arm_reshape_f32(const float32_t *input, float32_t *output, uint32_t total_size)
```

Reshape by copying data without changing element order.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input tensor data. |
| output | float32_t * | out | Pointer to the output tensor data. Nothing is copied when it aliases `input`. |
| total_size | uint32_t | in | Number of elements to copy. |

Source: `Include/arm_nnfunctions_flt.h:1404`

## arm_transpose_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_transpose_f16(
    const cmsis_nn_context *ctx,
    const cmsis_nn_transpose_params_f16 *params,
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *output_dims,
    float16_t *output
)
```

Transpose a floating-point tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in, out | Function context that may hold a temporary scratch buffer. |
| params | const cmsis_nn_transpose_params_f16 * | in | Transpose parameters, including permutation and layout information. num_dims must be in [1, 4] and perm must be a bijection over [0, num_dims - 1]. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. |
| input | const float16_t * | in | Pointer to the input tensor data. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. The first params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[perm[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not. |
| output | float16_t * | out | Pointer to the output tensor data. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3161`

## arm_concatenation_f16_x

`function` · `c`

```c
void arm_concatenation_f16_x(
    const float16_t *input,
    int32_t input_x,
    int32_t input_y,
    int32_t input_z,
    int32_t input_w,
    float16_t *output,
    int32_t output_x,
    uint32_t offset_x
)
```

Concatenate tensors along the X axis.

Call once per input tensor: `offset_x` selects where the input is stored along the X axis of the output tensor and must be advanced by `input_x` after each call. The output tensor must have the same height, channels and batch size as every input tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float16_t * | out | Pointer to the output tensor. |
| output_x | int32_t | in | Width of the output tensor. |
| offset_x | uint32_t | in | Offset on the X axis at which the input tensor is stored. Must be less than `output_x`. |

Source: `Include/arm_nnfunctions_flt.h:3171`

## arm_concatenation_f16_y

`function` · `c`

```c
void arm_concatenation_f16_y(
    const float16_t *input,
    int32_t input_x,
    int32_t input_y,
    int32_t input_z,
    int32_t input_w,
    float16_t *output,
    int32_t output_y,
    uint32_t offset_y
)
```

Concatenate tensors along the Y axis.

Call once per input tensor: `offset_y` selects where the input is stored along the Y axis of the output tensor and must be advanced by `input_y` after each call. The output tensor must have the same width, channels and batch size as every input tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float16_t * | out | Pointer to the output tensor. |
| output_y | int32_t | in | Height of the output tensor. |
| offset_y | uint32_t | in | Offset on the Y axis at which the input tensor is stored. Must be less than `output_y`. |

Source: `Include/arm_nnfunctions_flt.h:3183`

## arm_concatenation_f16_z

`function` · `c`

```c
void arm_concatenation_f16_z(
    const float16_t *input,
    int32_t input_x,
    int32_t input_y,
    int32_t input_z,
    int32_t input_w,
    float16_t *output,
    int32_t output_z,
    uint32_t offset_z
)
```

Concatenate tensors along the Z axis.

Call once per input tensor: `offset_z` selects where the input is stored along the Z axis of the output tensor and must be advanced by `input_z` after each call. The output tensor must have the same width, height and batch size as every input tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float16_t * | out | Pointer to the output tensor. |
| output_z | int32_t | in | Channels in the output tensor. |
| offset_z | uint32_t | in | Offset on the Z axis at which the input tensor is stored. Must be less than `output_z`. |

Source: `Include/arm_nnfunctions_flt.h:3195`

## arm_concatenation_f16_w

`function` · `c`

```c
void arm_concatenation_f16_w(
    const float16_t *input,
    int32_t input_x,
    int32_t input_y,
    int32_t input_z,
    int32_t input_w,
    float16_t *output,
    uint32_t offset_w
)
```

Concatenate tensors along the W axis.

Call once per input tensor: `offset_w` selects where the input is stored along the W axis of the output tensor and must be advanced by `input_w` after each call. The output tensor must have the same width, height and channels as every input tensor.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. Must not overlap the output tensor. |
| input_x | int32_t | in | Width of the input tensor. |
| input_y | int32_t | in | Height of the input tensor. |
| input_z | int32_t | in | Channels in the input tensor. |
| input_w | int32_t | in | Batch size in the input tensor. |
| output | float16_t * | out | Pointer to the output tensor. |
| offset_w | uint32_t | in | Offset on the W axis at which the input tensor is stored. |

Source: `Include/arm_nnfunctions_flt.h:3207`

## arm_concatenation_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_concatenation_f16(
    const float16_t *const *input_data,
    int32_t num_inputs,
    const int32_t *axis_sizes,
    int32_t output_dims,
    const int32_t *output_shape,
    int32_t axis,
    float16_t *output_data
)
```

Concatenate float32 tensors of any rank along one axis.

Rank-agnostic sibling of the 4-D per-axis arm_concatenation_f32_{x,y,z,w} entry points: all inputs at once, any rank, any axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Input `s` has the output shape with `output_shape`[axis] replaced by `axis_sizes`[s]; the inputs are laid down in order along the axis. Inputs must not overlap the output. A dimension of 0 is accepted and copies nothing.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t *const * | in | Array of `num_inputs` pointers to the flattened (row-major) inputs. |
| num_inputs | int32_t | in | Number of inputs (>= 1). |
| axis_sizes | const int32_t * | in | Array of length `num_inputs:` each input's extent along `axis`. |
| output_dims | int32_t | in | Number of dimensions in `output_shape` (>= 1). |
| output_shape | const int32_t * | in | Output shape; `output_shape`[axis] must equal the sum of `axis_sizes`. |
| axis | int32_t | in | Axis to concatenate along (0 <= axis < output_dims). |
| output_data | float16_t * | out | Pointer to the flattened output. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, size entry, size sum, NULL pointer or an element count above INT32_MAX. |

Source: `Include/arm_nnfunctions_flt.h:3218`

## arm_pack_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_pack_f16(
    const float16_t *const *input_data,
    int32_t num_inputs,
    int32_t input_dims,
    const int32_t *input_shape,
    int32_t axis,
    float16_t *output_data
)
```

Stack float32 tensors of equal shape along a new axis (TFLite PACK).

The output shape is `input_shape` with `num_inputs` inserted at `axis`; input `s` lands at index `s` of that axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Inputs must not overlap the output. Rank-0 inputs (`input_dims` == 0, `axis` == 0) stack into a vector.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t *const * | in | Array of `num_inputs` pointers to the flattened (row-major) inputs. |
| num_inputs | int32_t | in | Number of inputs (>= 1). |
| input_dims | int32_t | in | Number of dimensions of each input (>= 0). |
| input_shape | const int32_t * | in | Shape shared by every input (may be NULL when `input_dims` is 0). |
| axis | int32_t | in | Position of the new axis in the output (0 <= axis <= input_dims). |
| output_data | float16_t * | out | Pointer to the flattened output. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, NULL pointer or an element count above INT32_MAX. |

Source: `Include/arm_nnfunctions_flt.h:3229`

## arm_unpack_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_unpack_f16(
    const float16_t *input_data,
    int32_t input_dims,
    const int32_t *input_shape,
    int32_t axis,
    float16_t *const *output_data
)
```

Unstack a float32 tensor along one axis into `input_shape`[axis] tensors (TFLite UNPACK).

Inverse of arm_pack_f32: output `s` is the input with the axis fixed at index `s` and removed from the shape. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to the flattened (row-major) input. |
| input_dims | int32_t | in | Number of dimensions in `input_shape` (>= 1). |
| input_shape | const int32_t * | in | Input shape; `input_shape`[axis] (>= 1) is the number of outputs. |
| axis | int32_t | in | Axis to unstack (0 <= axis < input_dims). |
| output_data | float16_t *const * | out | Array of `input_shape`[axis] pointers to the flattened outputs. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, a zero-extent unstack axis (no outputs to produce), NULL pointer or an element count above INT32_MAX. |

Source: `Include/arm_nnfunctions_flt.h:3239`

## arm_batch_norm_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_batch_norm_f16(
    const float16_t *input,
    float16_t *output,
    const float16_t *scale,
    const float16_t *bias,
    const cmsis_nn_dims *input_dims,
    arm_nn_tensor_layout layout
)
```

Apply batch normalization.

Computes `output = input * scale[c] + bias[c]` for every element of channel `c`, with `scale` and `bias` holding the pre-folded per-channel factors.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor data. Format: [N, H, W, C]. |
| output | float16_t * | out | Pointer to the output tensor data, same shape as `input`. |
| scale | const float16_t * | in | Per-channel scale, `input_dims->c` values. |
| bias | const float16_t * | in | Per-channel bias, `input_dims->c` values. |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Every dimension must be positive. |
| layout | arm_nn_tensor_layout | in | Tensor layout selector. Must be `ARM_NN_LAYOUT_NHWC`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3272`

## arm_reshape_f16

`function` · `c`

```c
void arm_reshape_f16(const float16_t *input, float16_t *output, uint32_t total_size)
```

Reshape by copying data without changing element order.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor data. |
| output | float16_t * | out | Pointer to the output tensor data. Nothing is copied when it aliases `input`. |
| total_size | uint32_t | in | Number of elements to copy. |

Source: `Include/arm_nnfunctions_flt.h:3282`
