# heliaCORE.Reduction

## arm_reduce_sum_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_reduce_sum_f32(
    const float32_t *input_data,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *axis_dims,
    float32_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Computes the sum of the input tensor along the specified axes.

Sums are accumulated in float32 (also for the float16 variant, which rounds once to float16 at the end), so results do not overflow at float16 range and precision does not degrade with the reduction count. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (4D NHWC) |
| axis_dims | const cmsis_nn_dims * | in | 4D binary axis mask (non-zero = reduce that axis) |
| output_data | float32_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (reduced axes have size 1) |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1902`

## arm_argmin_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_argmin_f32(
    const float32_t *input_data,
    const cmsis_nn_dims *input_dims,
    int32_t axis,
    int32_t *output_data
)
```

Returns the first minimum's axis-relative INT32 index for a f32 tensor.

The input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.

A NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.

Metadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis | int32_t | in | Canonical reduction axis, in [0,3]. |
| output_data | int32_t * | out | INT32 indices, each in [0,input_dims[axis]). |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |

Source: `Include/arm_nnfunctions_flt.h:1940`

## arm_argmax_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_argmax_f32(
    const float32_t *input_data,
    const cmsis_nn_dims *input_dims,
    int32_t axis,
    int32_t *output_data
)
```

Returns the first maximum's axis-relative INT32 index for a f32 tensor.

The input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.

A NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.

Metadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis | int32_t | in | Canonical reduction axis, in [0,3]. |
| output_data | int32_t * | out | INT32 indices, each in [0,input_dims[axis]). |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |

Source: `Include/arm_nnfunctions_flt.h:1974`

## arm_reduce_max_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_reduce_max_f32(
    const float32_t *input_data,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *axis_dims,
    float32_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Reduces a f32 NHWC tensor to its maximum along a binary axis mask.

Values are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7fc00000); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.

A zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces -Inf; an empty output performs no accesses to data buffers.

All metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis_dims | const cmsis_nn_dims * | in | Four binary reduction flags. |
| output_data | float32_t * | out | Output tensor. |
| output_dims | const cmsis_nn_dims * | in | NHWC output shape with reduced axes retained as 1. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write. |

Source: `Include/arm_nnfunctions_flt.h:2004`

## arm_reduce_min_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_reduce_min_f32(
    const float32_t *input_data,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *axis_dims,
    float32_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Reduces a f32 NHWC tensor to its minimum along a binary axis mask.

Values are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7fc00000); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.

A zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces +Inf; an empty output performs no accesses to data buffers.

All metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis_dims | const cmsis_nn_dims * | in | Four binary reduction flags. |
| output_data | float32_t * | out | Output tensor. |
| output_dims | const cmsis_nn_dims * | in | NHWC output shape with reduced axes retained as 1. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write. |

Source: `Include/arm_nnfunctions_flt.h:2038`

## arm_nn_mean_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mean_f32(
    const float32_t *input_data,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *axis_dims,
    float32_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Computes the mean of a float32 tensor along the specified axes.

Values are accumulated and divided once in float32; unlike the float16 variant there is no wider accumulator, so rounding error can grow with the reduction length, matching arm_reduce_sum_f32. Because the intermediate accumulation is itself float32, it can saturate to +/-Inf even when the mean itself is representable, but whether it does depends on accumulation order: a strictly sequential build keeps one running sum, while vector builds  MVE intrinsics, or compiler auto-vectorization of the scalar path at -Ofast  fold per-lane partial sums, so on inputs whose partial sums exceed FLT_MAX in magnitude either build may return +/-Inf and the two may disagree (one finite, one Inf); when partial sums of opposite sign both saturate, the vector fold can even yield NaN (Inf + -Inf) from all-finite inputs. Only when every accumulation order overflows  e.g. same-signed values summing past FLT_MAX  is +/-Inf guaranteed on all builds. This is the accumulation-order divergence described below taken to the extreme. NaN and Inf propagate. A mean over all -0.0f inputs returns +0.0f on every build: the accumulator starts at +0.0f and (+0.0f) + (-0.0f) is +0.0f under round-to-nearest. Vector and scalar builds may differ in final ulps because float accumulation order differs.

Unlike arm_reduce_sum_f32 (identical signature, null checks only), this kernel validates shapes and returns `ARM_CMSIS_NN_ARG_ERROR` when any input dimension is less than 1, when any `output_dims` entry differs from the input shape with the reduced axes collapsed to 1, or when the input element count or the reduction count does not fit in int32_t. `output_data` must not overlap `input_data:` each output element is written after reading its whole reduction set, so an aliased write can corrupt inputs still to be read.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float32_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (4D NHWC) |
| axis_dims | const cmsis_nn_dims * | in | 4D binary axis mask (non-zero = reduce that axis) |
| output_data | float32_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (reduced axes have size 1) |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2086`

## arm_reduce_sum_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_reduce_sum_f16(
    const float16_t *input_data,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *axis_dims,
    float16_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Computes the sum of the input tensor along the specified axes.

Sums are accumulated in float32 (also for the float16 variant, which rounds once to float16 at the end), so results do not overflow at float16 range and precision does not degrade with the reduction count. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (4D NHWC) |
| axis_dims | const cmsis_nn_dims * | in | 4D binary axis mask (non-zero = reduce that axis) |
| output_data | float16_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (reduced axes have size 1) |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3646`

## arm_argmin_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_argmin_f16(
    const float16_t *input_data,
    const cmsis_nn_dims *input_dims,
    int32_t axis,
    int32_t *output_data
)
```

Returns the first minimum's axis-relative INT32 index for a f16 tensor.

The input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.

A NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.

Metadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis | int32_t | in | Canonical reduction axis, in [0,3]. |
| output_data | int32_t * | out | INT32 indices, each in [0,input_dims[axis]). |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |

Source: `Include/arm_nnfunctions_flt.h:3684`

## arm_argmax_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_argmax_f16(
    const float16_t *input_data,
    const cmsis_nn_dims *input_dims,
    int32_t axis,
    int32_t *output_data
)
```

Returns the first maximum's axis-relative INT32 index for a f16 tensor.

The input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.

A NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.

Metadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis | int32_t | in | Canonical reduction axis, in [0,3]. |
| output_data | int32_t * | out | INT32 indices, each in [0,input_dims[axis]). |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |

Source: `Include/arm_nnfunctions_flt.h:3718`

## arm_reduce_max_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_reduce_max_f16(
    const float16_t *input_data,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *axis_dims,
    float16_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Reduces a f16 NHWC tensor to its maximum along a binary axis mask.

Values are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7e00); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.

A zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces -Inf; an empty output performs no accesses to data buffers.

All metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis_dims | const cmsis_nn_dims * | in | Four binary reduction flags. |
| output_data | float16_t * | out | Output tensor. |
| output_dims | const cmsis_nn_dims * | in | NHWC output shape with reduced axes retained as 1. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write. |

Source: `Include/arm_nnfunctions_flt.h:3748`

## arm_reduce_min_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_reduce_min_f16(
    const float16_t *input_data,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *axis_dims,
    float16_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Reduces a f16 NHWC tensor to its minimum along a binary axis mask.

Values are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7e00); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.

A zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces +Inf; an empty output performs no accesses to data buffers.

All metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Input tensor. |
| input_dims | const cmsis_nn_dims * | in | Four NHWC extents. |
| axis_dims | const cmsis_nn_dims * | in | Four binary reduction flags. |
| output_data | float16_t * | out | Output tensor. |
| output_dims | const cmsis_nn_dims * | in | NHWC output shape with reduced axes retained as 1. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write. |

Source: `Include/arm_nnfunctions_flt.h:3782`

## arm_nn_mean_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_mean_f16(
    const float16_t *input_data,
    const cmsis_nn_dims *input_dims,
    const cmsis_nn_dims *axis_dims,
    float16_t *output_data,
    const cmsis_nn_dims *output_dims
)
```

Computes the mean of a float16 tensor along the specified axes.

Values are accumulated and divided in float32, then rounded once to float16. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs. Builds at -Ofast (the shipped CMSIS_OPTIMIZATION_LEVEL) may additionally differ from lower optimization levels by 1 ulp for non-power-of-two reduction counts: -freciprocal-math turns the divide-by-count into a multiply-by-reciprocal, which rounds differently.

Unlike arm_reduce_sum_f16 (identical signature, null checks only), this kernel validates shapes and returns `ARM_CMSIS_NN_ARG_ERROR` when any input dimension is less than 1, when any `output_dims` entry differs from the input shape with the reduced axes collapsed to 1, or when the input element count or the reduction count does not fit in int32_t. `output_data` must not overlap `input_data:` each output element is written after reading its whole reduction set, so an aliased write can corrupt inputs still to be read.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_data | const float16_t * | in | Pointer to input tensor |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (4D NHWC) |
| axis_dims | const cmsis_nn_dims * | in | 4D binary axis mask (non-zero = reduce that axis) |
| output_data | float16_t * | out | Pointer to output tensor |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (reduced axes have size 1) |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3817`
