Computes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.
Parameters
Parameters of arm_prelu_f32
Name
Type
Direction
Description
input_dims
const cmsis_nn_dims *
in
Input tensor dimensions. Must equal output_dims.
input
const float32_t *
in
Pointer to the input tensor.
alpha_dims
const cmsis_nn_dims *
in
Alpha tensor dimensions.
alpha
const float32_t *
in
Pointer to the alpha (slope) tensor.
output_dims
const cmsis_nn_dims *
in
Output tensor dimensions.
output
float32_t *
out
Pointer to the output tensor.
Returns
Returns of arm_prelu_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
Computes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise, evaluated as x * clamp(fma(x, 1/6, 0.5), 0, 1) so the saturated regions are exact: x >= 3 returns x bit-exactly and x <= -3 returns zero exactly (a negative zero, as IEEE negative * +0.0). In the curved region -3 < x < 3 the gate is a correctly rounded fused multiply-add on both build paths, so the scalar and MVE (cortex-m55) legs agree bit-exactly on every numeric normal input. Two carve-outs, both rooted in Armv8.1-M MVE floating-point arithmetic using the architecture’s Standard FPSCR value DN=1 and FZ=1 hard-wired, FZ16 passed through (Arm v8-M ARM, DDI 0553B.l, StandardFPSCRValue(), selected by the MVE FP pseudocode’s fpscr_controlled=FALSE): NaN lanes agree in NaN-ness but not necessarily in payload (forced DN makes the MVE leg canonicalize payloads the scalar leg preserves), and the MVE leg flushes f32 subnormal operands and results to a signed zero regardless of FPSCR.FZ, where the scalar leg with FZ clear keeps them. Both reference models (FVP Corstone-300 and QEMU mps3-an547) exhibit the flush identically; it has not been executed on silicon, where the same architectural behavior is required. Near the lower knot the absolute contract is the meaningful one: for x just above -3 the output error is dominated by the gate constant’s representation error, bounded by |x^2 * (1/6f - 1/6)| ~ 4.5e-08 near x = -3 (e.g. nextafterf(-3, 0) returns -7.45e-08 against a float64 -1.19e-07 millions of ulps of the tiny result, well inside the 1e-6 absolute contract), and where the gate underflows to exactly zero the kernel returns -0.0 with unbounded relative error. In-place operation (output == input) is supported on both legs; each element is read before it is written.
Parameters
Parameters of arm_hard_swish_f32
Name
Type
Direction
Description
input
const float32_t *
in
Pointer to the input samples.
output
float32_t *
out
Pointer to the output samples.
size
int32_t
in
Number of elements to process. Must be at least 1.
Returns
Returns of arm_hard_swish_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
Computes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.
Parameters
Parameters of arm_prelu_f16
Name
Type
Direction
Description
input_dims
const cmsis_nn_dims *
in
Input tensor dimensions. Must equal output_dims.
input
const float16_t *
in
Pointer to the input tensor.
alpha_dims
const cmsis_nn_dims *
in
Alpha tensor dimensions.
alpha
const float16_t *
in
Pointer to the alpha (slope) tensor.
output_dims
const cmsis_nn_dims *
in
Output tensor dimensions.
output
float16_t *
out
Pointer to the output tensor.
Returns
Returns of arm_prelu_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
Computes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise. The scalar leg widens each element to float32, evaluates the gate and the product there exactly as in arm_hard_swish_f32, and narrows only the final product, so it is single-rounded. The MVE (cortex-m55) leg evaluates the same expression in float16 throughout, scaling the gate by 1/6 before the product so that the multiplier stays in [0, 1]; it rounds the gate and the product separately and so can sit up to 2 float16 ulp away from the scalar leg in the curved region -3 < x < 3. The saturated regions are exact and identical on both legs (x >= 3 returns x bit-exactly, x <= -3 returns zero), as is the NaN/Inf behavior below; NaN lanes agree in NaN-ness but not necessarily in payload. In-place operation (output == input) is supported on both legs.
Parameters
Parameters of arm_hard_swish_f16
Name
Type
Direction
Description
input
const float16_t *
in
Pointer to the input samples.
output
float16_t *
out
Pointer to the output samples.
size
int32_t
in
Number of elements to process. Must be at least 1.
Returns
Returns of arm_hard_swish_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.