Computes the sum of the input tensor along the specified axes.
arm_cmsis_nn_status arm_reduce_sum_f32( const float32_t *input_data, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *axis_dims, float32_t *output_data, const cmsis_nn_dims *output_dims)Computes the sum of the input tensor along the specified axes.
Sums are accumulated in float32 (also for the float16 variant, which rounds once to float16 at the end), so results do not overflow at float16 range and precision does not degrade with the reduction count. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const float32_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (4D NHWC) |
axis_dims | const cmsis_nn_dims * | in | 4D binary axis mask (non-zero = reduce that axis) |
output_data | float32_t * | out | Pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions (reduced axes have size 1) |
Returns
| Description |
|---|
| `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |