Skip to content
heliaCORE
API reference
HELIA HUB

Activation Functions

Perform activation layers, including ReLU (Rectified Linear Unit), sigmoid and tanh

Machine-readable model

function

Elementwise activation.

Include/arm_nnfunctions_flt.h:609

arm_cmsis_nn_status arm_nn_activation_f32(
const float32_t *input,
float32_t *output,
int32_t size,
arm_nn_activation_type_flt type,
float32_t act_param
)

Elementwise activation.

Parameters of arm_nn_activation_f32
NameTypeDirectionDescription
inputconst float32_t *inPointer to the input samples.
outputfloat32_t *outPointer to the output samples.
sizeint32_tinNumber of elements to process.
typearm_nn_activation_type_fltinActivation selector.
act_paramfloat32_tinExtra activation parameter. Used for parameterized activations such as leaky ReLU.
Returns of arm_nn_activation_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Parametric ReLU for float32 data.

Include/arm_nnfunctions_flt.h:631

arm_cmsis_nn_status arm_prelu_f32(
const cmsis_nn_dims *input_dims,
const float32_t *input,
const cmsis_nn_dims *alpha_dims,
const float32_t *alpha,
const cmsis_nn_dims *output_dims,
float32_t *output
)

Parametric ReLU for float32 data.

Computes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.

Parameters of arm_prelu_f32
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Must equal output_dims.
inputconst float32_t *inPointer to the input tensor.
alpha_dimsconst cmsis_nn_dims *inAlpha tensor dimensions.
alphaconst float32_t *inPointer to the alpha (slope) tensor.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions.
outputfloat32_t *outPointer to the output tensor.
Returns of arm_prelu_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Hard swish activation for float32 data.

Include/arm_nnfunctions_flt.h:680

arm_cmsis_nn_status arm_hard_swish_f32(const float32_t *input, float32_t *output, int32_t size)

Hard swish activation for float32 data.

Computes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise, evaluated as x * clamp(fma(x, 1/6, 0.5), 0, 1) so the saturated regions are exact: x >= 3 returns x bit-exactly and x <= -3 returns zero exactly (a negative zero, as IEEE negative * +0.0). In the curved region -3 < x < 3 the gate is a correctly rounded fused multiply-add on both build paths, so the scalar and MVE (cortex-m55) legs agree bit-exactly on every numeric normal input. Two carve-outs, both rooted in Armv8.1-M MVE floating-point arithmetic using the architecture’s Standard FPSCR value DN=1 and FZ=1 hard-wired, FZ16 passed through (Arm v8-M ARM, DDI 0553B.l, StandardFPSCRValue(), selected by the MVE FP pseudocode’s fpscr_controlled=FALSE): NaN lanes agree in NaN-ness but not necessarily in payload (forced DN makes the MVE leg canonicalize payloads the scalar leg preserves), and the MVE leg flushes f32 subnormal operands and results to a signed zero regardless of FPSCR.FZ, where the scalar leg with FZ clear keeps them. Both reference models (FVP Corstone-300 and QEMU mps3-an547) exhibit the flush identically; it has not been executed on silicon, where the same architectural behavior is required. Near the lower knot the absolute contract is the meaningful one: for x just above -3 the output error is dominated by the gate constant’s representation error, bounded by |x^2 * (1/6f - 1/6)| ~ 4.5e-08 near x = -3 (e.g. nextafterf(-3, 0) returns -7.45e-08 against a float64 -1.19e-07 millions of ulps of the tiny result, well inside the 1e-6 absolute contract), and where the gate underflows to exactly zero the kernel returns -0.0 with unbounded relative error. In-place operation (output == input) is supported on both legs; each element is read before it is written.

Parameters of arm_hard_swish_f32
NameTypeDirectionDescription
inputconst float32_t *inPointer to the input samples.
outputfloat32_t *outPointer to the output samples.
sizeint32_tinNumber of elements to process. Must be at least 1.
Returns of arm_hard_swish_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Elementwise activation.

Include/arm_nnfunctions_flt.h:2750

arm_cmsis_nn_status arm_nn_activation_f16(
const float16_t *input,
float16_t *output,
int32_t size,
arm_nn_activation_type_flt type,
float16_t act_param
)

Elementwise activation.

Parameters of arm_nn_activation_f16
NameTypeDirectionDescription
inputconst float16_t *inPointer to the input samples.
outputfloat16_t *outPointer to the output samples.
sizeint32_tinNumber of elements to process.
typearm_nn_activation_type_fltinActivation selector.
act_paramfloat16_tinExtra activation parameter. Used for parameterized activations such as leaky ReLU.
Returns of arm_nn_activation_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Parametric ReLU for float32 data.

Include/arm_nnfunctions_flt.h:2759

arm_cmsis_nn_status arm_prelu_f16(
const cmsis_nn_dims *input_dims,
const float16_t *input,
const cmsis_nn_dims *alpha_dims,
const float16_t *alpha,
const cmsis_nn_dims *output_dims,
float16_t *output
)

Parametric ReLU for float32 data.

Computes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.

Parameters of arm_prelu_f16
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Must equal output_dims.
inputconst float16_t *inPointer to the input tensor.
alpha_dimsconst cmsis_nn_dims *inAlpha tensor dimensions.
alphaconst float16_t *inPointer to the alpha (slope) tensor.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions.
outputfloat16_t *outPointer to the output tensor.
Returns of arm_prelu_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Hard swish activation for float16 data.

Include/arm_nnfunctions_flt.h:2803

arm_cmsis_nn_status arm_hard_swish_f16(const float16_t *input, float16_t *output, int32_t size)

Hard swish activation for float16 data.

Computes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise. The scalar leg widens each element to float32, evaluates the gate and the product there exactly as in arm_hard_swish_f32, and narrows only the final product, so it is single-rounded. The MVE (cortex-m55) leg evaluates the same expression in float16 throughout, scaling the gate by 1/6 before the product so that the multiplier stays in [0, 1]; it rounds the gate and the product separately and so can sit up to 2 float16 ulp away from the scalar leg in the curved region -3 < x < 3. The saturated regions are exact and identical on both legs (x >= 3 returns x bit-exactly, x <= -3 returns zero), as is the NaN/Inf behavior below; NaN lanes agree in NaN-ness but not necessarily in payload. In-place operation (output == input) is supported on both legs.

Parameters of arm_hard_swish_f16
NameTypeDirectionDescription
inputconst float16_t *inPointer to the input samples.
outputfloat16_t *outPointer to the output samples.
sizeint32_tinNumber of elements to process. Must be at least 1.
Returns of arm_hard_swish_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.