# heliaCORE.groupElementwise

Elementwise add and multiplication functions.

## arm_elementwise_add_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_add_f32(
    const float32_t *input_1_vect,
    const float32_t *input_2_vect,
    float32_t *output,
    float32_t out_activation_min,
    float32_t out_activation_max,
    int32_t block_size
)
```

Elementwise add with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float32_t * | in | Pointer to the first input vector. |
| input_2_vect | const float32_t * | in | Pointer to the second input vector. |
| output | float32_t * | out | Pointer to the output vector. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:714`

## arm_elementwise_sub_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_sub_f32(
    const float32_t *input_1_vect,
    const float32_t *input_2_vect,
    float32_t *output,
    float32_t out_activation_min,
    float32_t out_activation_max,
    int32_t block_size
)
```

Elementwise subtract with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float32_t * | in | Pointer to the first input vector (minuend). |
| input_2_vect | const float32_t * | in | Pointer to the second input vector (subtrahend). |
| output | float32_t * | out | Pointer to the output vector. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:746`

## arm_nn_abs_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_abs_f32(const float32_t *input, float32_t *output, int32_t block_size)
```

Elementwise absolute value.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input vector. |
| output | float32_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:762`

## arm_nn_fill_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_fill_f32(float32_t value, float32_t *output, int32_t block_size)
```

Fill a float32 vector with one value.

Bit copy of `value` into every element (vector splat / plain stores), so a NaN fill value lands bit-exact, sign and payload included. Not named arm_fill_f32: CMSIS-DSP exports that symbol.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| value | float32_t | in | Fill value. |
| output | float32_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to write (0 is a no-op). |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`. |

Source: `Include/arm_nnfunctions_flt.h:778`

## arm_elementwise_mul_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_mul_f32(
    const float32_t *input_1_vect,
    const float32_t *input_2_vect,
    float32_t *output,
    float32_t out_activation_min,
    float32_t out_activation_max,
    int32_t block_size
)
```

Elementwise multiply with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float32_t * | in | Pointer to the first input vector. |
| input_2_vect | const float32_t * | in | Pointer to the second input vector. |
| output | float32_t * | out | Pointer to the output vector. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:805`

## arm_minimum_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_minimum_f32(
    const cmsis_nn_context *ctx,
    const float32_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float32_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float32_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.

The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float32_t * | in | Input 1, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Input 2, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |

Source: `Include/arm_nnfunctions_flt.h:858`

## arm_maximum_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_maximum_f32(
    const cmsis_nn_context *ctx,
    const float32_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float32_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float32_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.

The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float32_t * | in | Input 1, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Input 2, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |

Source: `Include/arm_nnfunctions_flt.h:894`

## arm_elementwise_sub_broadcast_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_sub_broadcast_f32(
    const float32_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float32_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float32_t *output_data,
    const cmsis_nn_dims *output_dims,
    float32_t out_activation_min,
    float32_t out_activation_max
)
```

Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and `output_dims` must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.

Numerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged  a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as `scalar - element`, not as the negation of `element - scalar`, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float32_t * | in | Minuend, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Subtrahend, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error. |

Source: `Include/arm_nnfunctions_flt.h:933`

## arm_elementwise_add_broadcast_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_add_broadcast_f32(
    const float32_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float32_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float32_t *output_data,
    const cmsis_nn_dims *output_dims,
    float32_t out_activation_min,
    float32_t out_activation_max
)
```

Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float32_t * | in | First input, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Second input, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |

Source: `Include/arm_nnfunctions_flt.h:957`

## arm_elementwise_mul_broadcast_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_mul_broadcast_f32(
    const float32_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float32_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float32_t *output_data,
    const cmsis_nn_dims *output_dims,
    float32_t out_activation_min,
    float32_t out_activation_max
)
```

Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float32_t * | in | First input, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float32_t * | in | Second input, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float32_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float32_t | in | Minimum output clamp value. |
| out_activation_max | float32_t | in | Maximum output clamp value. |

Source: `Include/arm_nnfunctions_flt.h:981`

## arm_nn_sqrt_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_sqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)
```

Elementwise square root.

The value path is scalar on every toolchain, because Helium has no vector square root. armclang and ATfE do vectorize the surrounding special-value classification; the results are bit-identical to the GCC scalar build, verified by executing both toolchains' objects (#295). Normal positive inputs evaluate `sqrtf(x)`, which IEEE 754 makes correctly rounded, so results are bit-exact to a float64 reference. Subnormal inputs follow FPSCR.FZ: where flush-to-zero is set - the Corstone-300 FVP default, and any host binary linked at -Ofast, where crtfastmath sets DAZ and FTZ - a subnormal input reads as zero and the result is +0. The float16 pair is immune, because it widens to a normal float32 first. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input vector. |
| output | float32_t * | out | Pointer to the output vector; may alias `input`. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1012`

## arm_rsqrt_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_rsqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)
```

Elementwise reciprocal square root, `1 / sqrt(x)`.

Same value path as arm_nn_sqrt_f32, including its FPSCR.FZ behaviour on subnormal inputs, where the result is +Inf. Normal positive inputs evaluate `1.0f / sqrtf(x)` in float32: two IEEE roundings, so the result is within 1 ulp of the correctly rounded value (measured against a float64 reference; `x = 4^k` is exact). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input vector. |
| output | float32_t * | out | Pointer to the output vector; may alias `input`. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1031`

## arm_elementwise_add_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_add_f16(
    const float16_t *input_1_vect,
    const float16_t *input_2_vect,
    float16_t *output,
    float16_t out_activation_min,
    float16_t out_activation_max,
    int32_t block_size
)
```

Elementwise add with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector. |
| input_2_vect | const float16_t * | in | Pointer to the second input vector. |
| output | float16_t * | out | Pointer to the output vector. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2815`

## arm_elementwise_add_fp16

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_add_fp16(
    const float16_t *input_1_vect,
    const float16_t *input_2_vect,
    float16_t *output,
    const float16_t out_activation_min,
    const float16_t out_activation_max,
    const int32_t block_size
)
```

Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate `arm_elementwise_add_f16()`. New code should call `arm_elementwise_add_f16()` instead.

This entry does NOT share the contract of `arm_elementwise_add_f16()`:

:::caution
No argument validation is performed. A NULL `input_1_vect`, `input_2_vect` or `output` is dereferenced rather than reported. A `block_size` of 0 writes nothing and still returns `ARM_CMSIS_NN_SUCCESS`, where `arm_elementwise_add_f16()` returns `ARM_CMSIS_NN_ARG_ERROR`.

:::

:::note
The clamp does not propagate NaN. Both the Helium path (`vminnm`/`vmaxnm`) and the scalar path (the non-propagating `MIN`/`MAX` clamp helper) bound against `out_activation_max` first, so a NaN produced by the addition comes back as `out_activation_max`. `arm_elementwise_add_f16()` documents TensorFlow Lite NaN propagation; this entry does not implement it.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector. Must not be NULL. |
| input_2_vect | const float16_t * | in | Pointer to the second input vector. Must not be NULL. |
| output | float16_t * | out | Pointer to the output vector. Must not be NULL. |
| out_activation_min | const float16_t | in | Minimum output clamp value. |
| out_activation_max | const float16_t | in | Maximum output clamp value. |
| block_size | const int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` unconditionally. |

Source: `Include/arm_nnfunctions_flt.h:2846`

## arm_elementwise_sub_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_sub_f16(
    const float16_t *input_1_vect,
    const float16_t *input_2_vect,
    float16_t *output,
    float16_t out_activation_min,
    float16_t out_activation_max,
    int32_t block_size
)
```

Elementwise subtract with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector (minuend). |
| input_2_vect | const float16_t * | in | Pointer to the second input vector (subtrahend). |
| output | float16_t * | out | Pointer to the output vector. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2856`

## arm_elementwise_squared_difference_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_squared_difference_f16(
    const float16_t *input_1_vect,
    const float16_t *input_2_vect,
    float16_t *output,
    int32_t block_size
)
```

Elementwise squared difference of two float16 vectors.

Each output element is calculated as `(input_1_vect[i] - input_2_vect[i])^2`.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector. |
| input_2_vect | const float16_t * | in | Pointer to the second input vector. |
| output | float16_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when an input/output pointer is NULL or `block_size` is less than 1. |

Source: `Include/arm_nnfunctions_flt.h:2876`

## arm_nn_abs_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_abs_f16(const float16_t *input, float16_t *output, int32_t block_size)
```

Elementwise absolute value.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input vector. |
| output | float16_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2884`

## arm_nn_fill_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_fill_f16(float16_t value, float16_t *output, int32_t block_size)
```

Fill a float16 vector with one value; bit copy of `value`, NaN payload included.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| value | float16_t | in | Fill value. |
| output | float16_t * | out | Pointer to the output vector. |
| block_size | int32_t | in | Number of elements to write (0 is a no-op). |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`. |

Source: `Include/arm_nnfunctions_flt.h:2896`

## arm_split_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_split_f16(
    const float16_t *input_data,
    const int32_t input_dims,
    const int32_t *input_shape,
    const int32_t axis,
    const int32_t num_splits,
    const int32_t *split_dims,
    float16_t *const *output_data
)
```

Split a float32 tensor of any rank into several tensors along one axis.

Inverse of arm_concatenation_f32; per-split lengths also cover SPLIT_V. Output `s` has the input shape with `input_shape`[axis] replaced by `split_dims`[s]. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input. A dimension of 0 is accepted and copies nothing.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to the flattened (row-major) input. |
| input_dims | const int32_t | in | Number of dimensions in `input_shape` (>= 1). |
| input_shape | const int32_t * | in | Input shape; `input_shape`[axis] must equal the sum of `split_dims`. |
| axis | const int32_t | in | Axis to split along (0 <= axis < input_dims). |
| num_splits | const int32_t | in | Number of outputs (>= 1). |
| split_dims | const int32_t * | in | Array of length `num_splits:` each output's extent along `axis`. |
| output_data | float16_t *const * | out | Array of `num_splits` pointers to the flattened outputs. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, split entry, split sum, NULL pointer or an element count above INT32_MAX. |

Source: `Include/arm_nnfunctions_flt.h:2923`

## arm_strided_slice_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_strided_slice_f16(
    const float16_t *input_data,
    float16_t *output_data,
    const cmsis_nn_dims *const input_dims,
    const cmsis_nn_dims *const begin_dims,
    const cmsis_nn_dims *const stride_dims,
    const cmsis_nn_dims *const output_dims
)
```

Strided slice for float32 data (pure copy, TensorFlow Lite compatible).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to input tensor. |
| output_data | float16_t * | out | Pointer to output tensor. |
| input_dims | const cmsis_nn_dims *const | in | Input tensor dimensions. |
| begin_dims | const cmsis_nn_dims *const | in | Begin dimensions for slicing. |
| stride_dims | const cmsis_nn_dims *const | in | Stride dimensions for slicing. |
| output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success. |

Source: `Include/arm_nnfunctions_flt.h:2934`

## arm_elementwise_mul_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_mul_f16(
    const float16_t *input_1_vect,
    const float16_t *input_2_vect,
    float16_t *output,
    float16_t out_activation_min,
    float16_t out_activation_max,
    int32_t block_size
)
```

Elementwise multiply with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY "no clamp" idiom).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_vect | const float16_t * | in | Pointer to the first input vector. |
| input_2_vect | const float16_t * | in | Pointer to the second input vector. |
| output | float16_t * | out | Pointer to the output vector. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |
| block_size | int32_t | in | Number of elements to process. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2944`

## arm_minimum_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_minimum_f16(
    const cmsis_nn_context *ctx,
    const float16_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float16_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float16_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.

The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float16_t * | in | Input 1, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Input 2, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |

Source: `Include/arm_nnfunctions_flt.h:2954`

## arm_maximum_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_maximum_f16(
    const cmsis_nn_context *ctx,
    const float16_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float16_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float16_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.

The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| ctx | const cmsis_nn_context * | in | Function context. Unused; may be NULL. |
| input_1_data | const float16_t * | in | Input 1, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Input 2, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |

Source: `Include/arm_nnfunctions_flt.h:2965`

## arm_elementwise_sub_broadcast_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_sub_broadcast_f16(
    const float16_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float16_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float16_t *output_data,
    const cmsis_nn_dims *output_dims,
    float16_t out_activation_min,
    float16_t out_activation_max
)
```

Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and `output_dims` must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.

Numerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged  a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as `scalar - element`, not as the negation of `element - scalar`, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.

Half-precision twin: the numerics are those of arm_elementwise_sub_f16 on the materialised operands.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float16_t * | in | Minuend, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Subtrahend, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error. |

Source: `Include/arm_nnfunctions_flt.h:2978`

## arm_elementwise_add_broadcast_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_add_broadcast_f16(
    const float16_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float16_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float16_t *output_data,
    const cmsis_nn_dims *output_dims,
    float16_t out_activation_min,
    float16_t out_activation_max
)
```

Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.

Half-precision twin: the numerics are those of arm_elementwise_add_f16 on the materialised operands.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float16_t * | in | First input, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Second input, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |

Source: `Include/arm_nnfunctions_flt.h:2992`

## arm_elementwise_mul_broadcast_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_elementwise_mul_broadcast_f16(
    const float16_t *input_1_data,
    const cmsis_nn_dims *input_1_dims,
    const float16_t *input_2_data,
    const cmsis_nn_dims *input_2_dims,
    float16_t *output_data,
    const cmsis_nn_dims *output_dims,
    float16_t out_activation_min,
    float16_t out_activation_max
)
```

Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.

Half-precision twin: the numerics are those of arm_elementwise_mul_f16 on the materialised operands.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_1_data | const float16_t * | in | First input, NHWC, sized by `input_1_dims`. |
| input_1_dims | const cmsis_nn_dims * | in | Dimensions of input 1. |
| input_2_data | const float16_t * | in | Second input, NHWC, sized by `input_2_dims`. |
| input_2_dims | const cmsis_nn_dims * | in | Dimensions of input 2. |
| output_data | float16_t * | out | Output, NHWC, sized by `output_dims`. |
| output_dims | const cmsis_nn_dims * | in | Broadcast output dimensions. |
| out_activation_min | float16_t | in | Minimum output clamp value. |
| out_activation_max | float16_t | in | Maximum output clamp value. |

Source: `Include/arm_nnfunctions_flt.h:3006`

## arm_nn_sqrt_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_sqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)
```

Elementwise square root of a float16 tensor.

The value path is scalar on every toolchain, because Helium has no vector square root; armclang and ATfE vectorize the surrounding classification into an MVE loop and produce bit-identical results, verified by executing their objects (#295). Each element is widened to float32, `sqrtf` is evaluated there and the result is rounded once to float16. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 square root (0 ulp, #295). Widening first also makes this pair immune to FPSCR.FZ, which flushes float32 subnormals in the f32 pair. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. |
| output | float16_t * | out | Pointer to the output tensor; may alias `input`. |
| block_size | int32_t | in | Number of tensor elements. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3036`

## arm_rsqrt_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_rsqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)
```

Elementwise reciprocal square root of a float16 tensor, `1 / sqrt(x)`.

Same value path as arm_nn_sqrt_f16: widen to float32, evaluate `1.0f / sqrtf(x)` there, round once to float16, and so also immune to FPSCR.FZ. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 reciprocal square root (0 ulp, #295). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input tensor. |
| output | float16_t * | out | Pointer to the output tensor; may alias `input`. |
| block_size | int32_t | in | Number of tensor elements. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3055`
