Elementwise add with optional output clamp.
arm_cmsis_nn_status arm_elementwise_add_f32( const float32_t *input_1_vect, const float32_t *input_2_vect, float32_t *output, float32_t out_activation_min, float32_t out_activation_max, int32_t block_size)Elementwise add with optional output clamp.
NaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined cross-toolchain on-target execution is #340’s scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY “no clamp” idiom).
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const float32_t * | in | Pointer to the first input vector. |
input_2_vect | const float32_t * | in | Pointer to the second input vector. |
output | float32_t * | out | Pointer to the output vector. |
out_activation_min | float32_t | in | Minimum output clamp value. |
out_activation_max | float32_t | in | Maximum output clamp value. |
block_size | int32_t | in | Number of elements to process. |
Returns
| Description |
|---|
| `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |