Skip to content
heliaCORE
API reference
HELIA HUB

Elementwise Functions

Elementwise add and multiplication functions.

Machine-readable model

function

Elementwise add with optional output clamp.

Include/arm_nnfunctions_flt.h:714

arm_cmsis_nn_status arm_elementwise_add_f32(
const float32_t *input_1_vect,
const float32_t *input_2_vect,
float32_t *output,
float32_t out_activation_min,
float32_t out_activation_max,
int32_t block_size
)

Elementwise add with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340’s scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY “no clamp” idiom).

Parameters of arm_elementwise_add_f32
NameTypeDirectionDescription
input_1_vectconst float32_t *inPointer to the first input vector.
input_2_vectconst float32_t *inPointer to the second input vector.
outputfloat32_t *outPointer to the output vector.
out_activation_minfloat32_tinMinimum output clamp value.
out_activation_maxfloat32_tinMaximum output clamp value.
block_sizeint32_tinNumber of elements to process.
Returns of arm_elementwise_add_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Elementwise subtract with optional output clamp.

Include/arm_nnfunctions_flt.h:746

arm_cmsis_nn_status arm_elementwise_sub_f32(
const float32_t *input_1_vect,
const float32_t *input_2_vect,
float32_t *output,
float32_t out_activation_min,
float32_t out_activation_max,
int32_t block_size
)

Elementwise subtract with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340’s scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY “no clamp” idiom).

Parameters of arm_elementwise_sub_f32
NameTypeDirectionDescription
input_1_vectconst float32_t *inPointer to the first input vector (minuend).
input_2_vectconst float32_t *inPointer to the second input vector (subtrahend).
outputfloat32_t *outPointer to the output vector.
out_activation_minfloat32_tinMinimum output clamp value.
out_activation_maxfloat32_tinMaximum output clamp value.
block_sizeint32_tinNumber of elements to process.
Returns of arm_elementwise_sub_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Elementwise absolute value.

Include/arm_nnfunctions_flt.h:762

arm_cmsis_nn_status arm_nn_abs_f32(const float32_t *input, float32_t *output, int32_t block_size)

Elementwise absolute value.

Parameters of arm_nn_abs_f32
NameTypeDirectionDescription
inputconst float32_t *inPointer to the input vector.
outputfloat32_t *outPointer to the output vector.
block_sizeint32_tinNumber of elements to process.
Returns of arm_nn_abs_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Fill a float32 vector with one value.

Include/arm_nnfunctions_flt.h:778

arm_cmsis_nn_status arm_nn_fill_f32(float32_t value, float32_t *output, int32_t block_size)

Fill a float32 vector with one value.

Bit copy of value into every element (vector splat / plain stores), so a NaN fill value lands bit-exact, sign and payload included. Not named arm_fill_f32: CMSIS-DSP exports that symbol.

Parameters of arm_nn_fill_f32
NameTypeDirectionDescription
valuefloat32_tinFill value.
outputfloat32_t *outPointer to the output vector.
block_sizeint32_tinNumber of elements to write (0 is a no-op).
Returns of arm_nn_fill_f32
Description
`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`.
function

Elementwise multiply with optional output clamp.

Include/arm_nnfunctions_flt.h:805

arm_cmsis_nn_status arm_elementwise_mul_f32(
const float32_t *input_1_vect,
const float32_t *input_2_vect,
float32_t *output,
float32_t out_activation_min,
float32_t out_activation_max,
int32_t block_size
)

Elementwise multiply with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340’s scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY “no clamp” idiom).

Parameters of arm_elementwise_mul_f32
NameTypeDirectionDescription
input_1_vectconst float32_t *inPointer to the first input vector.
input_2_vectconst float32_t *inPointer to the second input vector.
outputfloat32_t *outPointer to the output vector.
out_activation_minfloat32_tinMinimum output clamp value.
out_activation_maxfloat32_tinMaximum output clamp value.
block_sizeint32_tinNumber of elements to process.
Returns of arm_elementwise_mul_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape.

Include/arm_nnfunctions_flt.h:858

arm_cmsis_nn_status arm_minimum_f32(
const cmsis_nn_context *ctx,
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)

Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and output_dims must be their broadcast shape.

The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.

Parameters of arm_minimum_f32
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. Unused; may be NULL.
input_1_dataconst float32_t *inInput 1, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float32_t *inInput 2, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat32_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
Returns of arm_minimum_f32
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape.

Include/arm_nnfunctions_flt.h:894

arm_cmsis_nn_status arm_maximum_f32(
const cmsis_nn_context *ctx,
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims
)

Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and output_dims must be their broadcast shape.

The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.

Parameters of arm_maximum_f32
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. Unused; may be NULL.
input_1_dataconst float32_t *inInput 1, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float32_t *inInput 2, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat32_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
Returns of arm_maximum_f32
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.

Include/arm_nnfunctions_flt.h:933

arm_cmsis_nn_status arm_elementwise_sub_broadcast_f32(
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims,
float32_t out_activation_min,
float32_t out_activation_max
)

Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and output_dims must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.

Numerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as scalar - element, not as the negation of element - scalar, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.

Parameters of arm_elementwise_sub_broadcast_f32
NameTypeDirectionDescription
input_1_dataconst float32_t *inMinuend, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float32_t *inSubtrahend, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat32_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
out_activation_minfloat32_tinMinimum output clamp value.
out_activation_maxfloat32_tinMaximum output clamp value.
Returns of arm_elementwise_sub_broadcast_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error.
function

Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.

Include/arm_nnfunctions_flt.h:957

arm_cmsis_nn_status arm_elementwise_add_broadcast_f32(
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims,
float32_t out_activation_min,
float32_t out_activation_max
)

Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.

Parameters of arm_elementwise_add_broadcast_f32
NameTypeDirectionDescription
input_1_dataconst float32_t *inFirst input, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float32_t *inSecond input, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat32_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
out_activation_minfloat32_tinMinimum output clamp value.
out_activation_maxfloat32_tinMaximum output clamp value.
function

Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.

Include/arm_nnfunctions_flt.h:981

arm_cmsis_nn_status arm_elementwise_mul_broadcast_f32(
const float32_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float32_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float32_t *output_data,
const cmsis_nn_dims *output_dims,
float32_t out_activation_min,
float32_t out_activation_max
)

Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.

Parameters of arm_elementwise_mul_broadcast_f32
NameTypeDirectionDescription
input_1_dataconst float32_t *inFirst input, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float32_t *inSecond input, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat32_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
out_activation_minfloat32_tinMinimum output clamp value.
out_activation_maxfloat32_tinMaximum output clamp value.
function

Elementwise square root.

Include/arm_nnfunctions_flt.h:1012

arm_cmsis_nn_status arm_nn_sqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)

Elementwise square root.

The value path is scalar on every toolchain, because Helium has no vector square root. armclang and ATfE do vectorize the surrounding special-value classification; the results are bit-identical to the GCC scalar build, verified by executing both toolchains’ objects (#295). Normal positive inputs evaluate sqrtf(x), which IEEE 754 makes correctly rounded, so results are bit-exact to a float64 reference. Subnormal inputs follow FPSCR.FZ: where flush-to-zero is set - the Corstone-300 FVP default, and any host binary linked at -Ofast, where crtfastmath sets DAZ and FTZ - a subnormal input reads as zero and the result is +0. The float16 pair is immune, because it widens to a normal float32 first. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).

Parameters of arm_nn_sqrt_f32
NameTypeDirectionDescription
inputconst float32_t *inPointer to the input vector.
outputfloat32_t *outPointer to the output vector; may alias `input`.
block_sizeint32_tinNumber of elements to process.
Returns of arm_nn_sqrt_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Elementwise reciprocal square root, 1 / sqrt(x).

Include/arm_nnfunctions_flt.h:1031

arm_cmsis_nn_status arm_rsqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)

Elementwise reciprocal square root, 1 / sqrt(x).

Same value path as arm_nn_sqrt_f32, including its FPSCR.FZ behaviour on subnormal inputs, where the result is +Inf. Normal positive inputs evaluate 1.0f / sqrtf(x) in float32: two IEEE roundings, so the result is within 1 ulp of the correctly rounded value (measured against a float64 reference; x = 4^k is exact). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).

Parameters of arm_rsqrt_f32
NameTypeDirectionDescription
inputconst float32_t *inPointer to the input vector.
outputfloat32_t *outPointer to the output vector; may alias `input`.
block_sizeint32_tinNumber of elements to process.
Returns of arm_rsqrt_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Elementwise add with optional output clamp.

Include/arm_nnfunctions_flt.h:2815

arm_cmsis_nn_status arm_elementwise_add_f16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
float16_t out_activation_min,
float16_t out_activation_max,
int32_t block_size
)

Elementwise add with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340’s scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY “no clamp” idiom).

Parameters of arm_elementwise_add_f16
NameTypeDirectionDescription
input_1_vectconst float16_t *inPointer to the first input vector.
input_2_vectconst float16_t *inPointer to the second input vector.
outputfloat16_t *outPointer to the output vector.
out_activation_minfloat16_tinMinimum output clamp value.
out_activation_maxfloat16_tinMaximum output clamp value.
block_sizeint32_tinNumber of elements to process.
Returns of arm_elementwise_add_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate armelementwiseaddf16().

Include/arm_nnfunctions_flt.h:2846

arm_cmsis_nn_status arm_elementwise_add_fp16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
const float16_t out_activation_min,
const float16_t out_activation_max,
const int32_t block_size
)

Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate arm_elementwise_add_f16(). New code should call arm_elementwise_add_f16() instead.

This entry does NOT share the contract of arm_elementwise_add_f16():

Parameters of arm_elementwise_add_fp16
NameTypeDirectionDescription
input_1_vectconst float16_t *inPointer to the first input vector. Must not be NULL.
input_2_vectconst float16_t *inPointer to the second input vector. Must not be NULL.
outputfloat16_t *outPointer to the output vector. Must not be NULL.
out_activation_minconst float16_tinMinimum output clamp value.
out_activation_maxconst float16_tinMaximum output clamp value.
block_sizeconst int32_tinNumber of elements to process.
Returns of arm_elementwise_add_fp16
Description
`ARM_CMSIS_NN_SUCCESS` unconditionally.
function

Elementwise subtract with optional output clamp.

Include/arm_nnfunctions_flt.h:2856

arm_cmsis_nn_status arm_elementwise_sub_f16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
float16_t out_activation_min,
float16_t out_activation_max,
int32_t block_size
)

Elementwise subtract with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340’s scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY “no clamp” idiom).

Parameters of arm_elementwise_sub_f16
NameTypeDirectionDescription
input_1_vectconst float16_t *inPointer to the first input vector (minuend).
input_2_vectconst float16_t *inPointer to the second input vector (subtrahend).
outputfloat16_t *outPointer to the output vector.
out_activation_minfloat16_tinMinimum output clamp value.
out_activation_maxfloat16_tinMaximum output clamp value.
block_sizeint32_tinNumber of elements to process.
Returns of arm_elementwise_sub_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Elementwise squared difference of two float16 vectors.

Include/arm_nnfunctions_flt.h:2876

arm_cmsis_nn_status arm_elementwise_squared_difference_f16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
int32_t block_size
)

Elementwise squared difference of two float16 vectors.

Each output element is calculated as (input_1_vect[i] - input_2_vect[i])^2.

Parameters of arm_elementwise_squared_difference_f16
NameTypeDirectionDescription
input_1_vectconst float16_t *inPointer to the first input vector.
input_2_vectconst float16_t *inPointer to the second input vector.
outputfloat16_t *outPointer to the output vector.
block_sizeint32_tinNumber of elements to process.
Returns of arm_elementwise_squared_difference_f16
Description
`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when an input/output pointer is NULL or `block_size` is less than 1.
function

Elementwise absolute value.

Include/arm_nnfunctions_flt.h:2884

arm_cmsis_nn_status arm_nn_abs_f16(const float16_t *input, float16_t *output, int32_t block_size)

Elementwise absolute value.

Parameters of arm_nn_abs_f16
NameTypeDirectionDescription
inputconst float16_t *inPointer to the input vector.
outputfloat16_t *outPointer to the output vector.
block_sizeint32_tinNumber of elements to process.
Returns of arm_nn_abs_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Fill a float16 vector with one value; bit copy of value, NaN payload included.

Include/arm_nnfunctions_flt.h:2896

arm_cmsis_nn_status arm_nn_fill_f16(float16_t value, float16_t *output, int32_t block_size)

Fill a float16 vector with one value; bit copy of value, NaN payload included.

Parameters of arm_nn_fill_f16
NameTypeDirectionDescription
valuefloat16_tinFill value.
outputfloat16_t *outPointer to the output vector.
block_sizeint32_tinNumber of elements to write (0 is a no-op).
Returns of arm_nn_fill_f16
Description
`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`.
function

Split a float32 tensor of any rank into several tensors along one axis.

Include/arm_nnfunctions_flt.h:2923

arm_cmsis_nn_status arm_split_f16(
const float16_t *input_data,
const int32_t input_dims,
const int32_t *input_shape,
const int32_t axis,
const int32_t num_splits,
const int32_t *split_dims,
float16_t *const *output_data
)

Split a float32 tensor of any rank into several tensors along one axis.

Inverse of arm_concatenation_f32; per-split lengths also cover SPLIT_V. Output s has the input shape with input_shape[axis] replaced by split_dims[s]. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input. A dimension of 0 is accepted and copies nothing.

Parameters of arm_split_f16
NameTypeDirectionDescription
input_dataconst float16_t *inPointer to the flattened (row-major) input.
input_dimsconst int32_tinNumber of dimensions in `input_shape` (>= 1).
input_shapeconst int32_t *inInput shape; `input_shape`[axis] must equal the sum of `split_dims`.
axisconst int32_tinAxis to split along (0 <= axis < input_dims).
num_splitsconst int32_tinNumber of outputs (>= 1).
split_dimsconst int32_t *inArray of length `num_splits:` each output's extent along `axis`.
output_datafloat16_t *const *outArray of `num_splits` pointers to the flattened outputs.
Returns of arm_split_f16
Description
`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, split entry, split sum, NULL pointer or an element count above INT32_MAX.
function

Strided slice for float32 data (pure copy, TensorFlow Lite compatible).

Include/arm_nnfunctions_flt.h:2934

arm_cmsis_nn_status arm_strided_slice_f16(
const float16_t *input_data,
float16_t *output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const begin_dims,
const cmsis_nn_dims *const stride_dims,
const cmsis_nn_dims *const output_dims
)

Strided slice for float32 data (pure copy, TensorFlow Lite compatible).

Parameters of arm_strided_slice_f16
NameTypeDirectionDescription
input_dataconst float16_t *inPointer to input tensor.
output_datafloat16_t *outPointer to output tensor.
input_dimsconst cmsis_nn_dims *constinInput tensor dimensions.
begin_dimsconst cmsis_nn_dims *constinBegin dimensions for slicing.
stride_dimsconst cmsis_nn_dims *constinStride dimensions for slicing.
output_dimsconst cmsis_nn_dims *constinOutput tensor dimensions.
Returns of arm_strided_slice_f16
Description
ARM_CMSIS_NN_SUCCESS on success.
function

Elementwise multiply with optional output clamp.

Include/arm_nnfunctions_flt.h:2944

arm_cmsis_nn_status arm_elementwise_mul_f16(
const float16_t *input_1_vect,
const float16_t *input_2_vect,
float16_t *output,
float16_t out_activation_min,
float16_t out_activation_max,
int32_t block_size
)

Elementwise multiply with optional output clamp.

NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340’s scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY “no clamp” idiom).

Parameters of arm_elementwise_mul_f16
NameTypeDirectionDescription
input_1_vectconst float16_t *inPointer to the first input vector.
input_2_vectconst float16_t *inPointer to the second input vector.
outputfloat16_t *outPointer to the output vector.
out_activation_minfloat16_tinMinimum output clamp value.
out_activation_maxfloat16_tinMaximum output clamp value.
block_sizeint32_tinNumber of elements to process.
Returns of arm_elementwise_mul_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape.

Include/arm_nnfunctions_flt.h:2954

arm_cmsis_nn_status arm_minimum_f16(
const cmsis_nn_context *ctx,
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)

Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and output_dims must be their broadcast shape.

The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.

Parameters of arm_minimum_f16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. Unused; may be NULL.
input_1_dataconst float16_t *inInput 1, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float16_t *inInput 2, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat16_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
Returns of arm_minimum_f16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape.

Include/arm_nnfunctions_flt.h:2965

arm_cmsis_nn_status arm_maximum_f16(
const cmsis_nn_context *ctx,
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims
)

Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and output_dims must be their broadcast shape.

The result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.

Parameters of arm_maximum_f16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. Unused; may be NULL.
input_1_dataconst float16_t *inInput 1, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float16_t *inInput 2, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat16_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
Returns of arm_maximum_f16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.

Include/arm_nnfunctions_flt.h:2978

arm_cmsis_nn_status arm_elementwise_sub_broadcast_f16(
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims,
float16_t out_activation_min,
float16_t out_activation_max
)

Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and output_dims must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.

Numerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as scalar - element, not as the negation of element - scalar, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.

Half-precision twin: the numerics are those of arm_elementwise_sub_f16 on the materialised operands.

Parameters of arm_elementwise_sub_broadcast_f16
NameTypeDirectionDescription
input_1_dataconst float16_t *inMinuend, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float16_t *inSubtrahend, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat16_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
out_activation_minfloat16_tinMinimum output clamp value.
out_activation_maxfloat16_tinMaximum output clamp value.
Returns of arm_elementwise_sub_broadcast_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error.
function

Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.

Include/arm_nnfunctions_flt.h:2992

arm_cmsis_nn_status arm_elementwise_add_broadcast_f16(
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims,
float16_t out_activation_min,
float16_t out_activation_max
)

Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.

Half-precision twin: the numerics are those of arm_elementwise_add_f16 on the materialised operands.

Parameters of arm_elementwise_add_broadcast_f16
NameTypeDirectionDescription
input_1_dataconst float16_t *inFirst input, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float16_t *inSecond input, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat16_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
out_activation_minfloat16_tinMinimum output clamp value.
out_activation_maxfloat16_tinMaximum output clamp value.
function

Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.

Include/arm_nnfunctions_flt.h:3006

arm_cmsis_nn_status arm_elementwise_mul_broadcast_f16(
const float16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const float16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
float16_t *output_data,
const cmsis_nn_dims *output_dims,
float16_t out_activation_min,
float16_t out_activation_max
)

Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.

Broadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.

Half-precision twin: the numerics are those of arm_elementwise_mul_f16 on the materialised operands.

Parameters of arm_elementwise_mul_broadcast_f16
NameTypeDirectionDescription
input_1_dataconst float16_t *inFirst input, NHWC, sized by `input_1_dims`.
input_1_dimsconst cmsis_nn_dims *inDimensions of input 1.
input_2_dataconst float16_t *inSecond input, NHWC, sized by `input_2_dims`.
input_2_dimsconst cmsis_nn_dims *inDimensions of input 2.
output_datafloat16_t *outOutput, NHWC, sized by `output_dims`.
output_dimsconst cmsis_nn_dims *inBroadcast output dimensions.
out_activation_minfloat16_tinMinimum output clamp value.
out_activation_maxfloat16_tinMaximum output clamp value.
function

Elementwise square root of a float16 tensor.

Include/arm_nnfunctions_flt.h:3036

arm_cmsis_nn_status arm_nn_sqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)

Elementwise square root of a float16 tensor.

The value path is scalar on every toolchain, because Helium has no vector square root; armclang and ATfE vectorize the surrounding classification into an MVE loop and produce bit-identical results, verified by executing their objects (#295). Each element is widened to float32, sqrtf is evaluated there and the result is rounded once to float16. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 square root (0 ulp, #295). Widening first also makes this pair immune to FPSCR.FZ, which flushes float32 subnormals in the f32 pair. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).

Parameters of arm_nn_sqrt_f16
NameTypeDirectionDescription
inputconst float16_t *inPointer to the input tensor.
outputfloat16_t *outPointer to the output tensor; may alias `input`.
block_sizeint32_tinNumber of tensor elements.
Returns of arm_nn_sqrt_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Elementwise reciprocal square root of a float16 tensor, 1 / sqrt(x).

Include/arm_nnfunctions_flt.h:3055

arm_cmsis_nn_status arm_rsqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)

Elementwise reciprocal square root of a float16 tensor, 1 / sqrt(x).

Same value path as arm_nn_sqrt_f16: widen to float32, evaluate 1.0f / sqrtf(x) there, round once to float16, and so also immune to FPSCR.FZ. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 reciprocal square root (0 ulp, #295). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).

Parameters of arm_rsqrt_f16
NameTypeDirectionDescription
inputconst float16_t *inPointer to the input tensor.
outputfloat16_t *outPointer to the output tensor; may alias `input`.
block_sizeint32_tinNumber of tensor elements.
Returns of arm_rsqrt_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.