# heliaCORE.Acti

Perform activation layers, including ReLU (Rectified Linear Unit), sigmoid and tanh

## arm_nn_activation_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_activation_f32(
    const float32_t *input,
    float32_t *output,
    int32_t size,
    arm_nn_activation_type_flt type,
    float32_t act_param
)
```

Elementwise activation.

:::note
The RELU, RELU6 and LEAKY_RELU legs classify NaN on the integer bit pattern (#380 / #382), so a NaN input comes back as NaN at every optimization level on the gated toolchains, including the shipped -Ofast. This holds on both the scalar and the MVE (cortex-m55) build paths; the MVE RELU/RELU6 legs restore the NaN lanes that vmaxnmq/vminnmq suppress. The MVE TANH leg returns a NaN input unchanged and keeps the sign of zero, also decided on the bit pattern (#635). The scalar TANH leg returns NaN where there is no hardware floating point (__ARM_FP undefined, e.g. Cortex-M0), where it too classifies NaN on the bit pattern (quieting a signalling NaN), and elsewhere only in builds without -ffinite-math-only. SIGMOID and HARDSWISH are outside this contract; see the per-helper notes in `Include/Internal/arm_nn_activation_flt.h`.

:::

:::note
The HARDSWISH leg's scalar helper (arm_nn_hardswish_scalar_f32, serving every build that does not take the MVE float path  no MVE float support, or MVE present but not used, e.g. under ARM_MATH_AUTOVECTORIZE) keeps the legacy separately rounded multiply-and-add gate and can differ by an ulp in the curved region from the standalone `arm_hard_swish_f32`, whose gate is a correctly rounded fma; the mux's MVE helper (arm_nn_vhardswish_mve_f32) uses vfmaq and agrees with that kernel. Callers that need bit-exact, leg-agreeing hard swish  or the documented NaN/Inf contract  should call `arm_hard_swish_f32` directly.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input samples. |
| output | float32_t * | out | Pointer to the output samples. |
| size | int32_t | in | Number of elements to process. |
| type | arm_nn_activation_type_flt | in | Activation selector. |
| act_param | float32_t | in | Extra activation parameter. Used for parameterized activations such as leaky ReLU. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:609`

## arm_prelu_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_prelu_f32(
    const cmsis_nn_dims *input_dims,
    const float32_t *input,
    const cmsis_nn_dims *alpha_dims,
    const float32_t *alpha,
    const cmsis_nn_dims *output_dims,
    float32_t *output
)
```

Parametric ReLU for float32 data.

Computes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Must equal output_dims. |
| input | const float32_t * | in | Pointer to the input tensor. |
| alpha_dims | const cmsis_nn_dims * | in | Alpha tensor dimensions. |
| alpha | const float32_t * | in | Pointer to the alpha (slope) tensor. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float32_t * | out | Pointer to the output tensor. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:631`

## arm_hard_swish_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_hard_swish_f32(const float32_t *input, float32_t *output, int32_t size)
```

Hard swish activation for float32 data.

Computes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise, evaluated as x * clamp(fma(x, 1/6, 0.5), 0, 1) so the saturated regions are exact: x >= 3 returns x bit-exactly and x <= -3 returns zero exactly (a negative zero, as IEEE negative * +0.0). In the curved region -3 < x < 3 the gate is a correctly rounded fused multiply-add on both build paths, so the scalar and MVE (cortex-m55) legs agree bit-exactly on every numeric normal input. Two carve-outs, both rooted in Armv8.1-M MVE floating-point arithmetic using the architecture's Standard FPSCR value  DN=1 and FZ=1 hard-wired, FZ16 passed through (Arm v8-M ARM, DDI 0553B.l, StandardFPSCRValue(), selected by the MVE FP pseudocode's fpscr_controlled=FALSE): NaN lanes agree in NaN-ness but not necessarily in payload (forced DN makes the MVE leg canonicalize payloads the scalar leg preserves), and the MVE leg flushes f32 subnormal operands and results to a signed zero regardless of FPSCR.FZ, where the scalar leg with FZ clear keeps them. Both reference models (FVP Corstone-300 and QEMU mps3-an547) exhibit the flush identically; it has not been executed on silicon, where the same architectural behavior is required. Near the lower knot the absolute contract is the meaningful one: for x just above -3 the output error is dominated by the gate constant's representation error, bounded by |x^2 * (1/6f - 1/6)| ~ 4.5e-08 near x = -3 (e.g. nextafterf(-3, 0) returns -7.45e-08 against a float64 -1.19e-07  millions of ulps of the tiny result, well inside the 1e-6 absolute contract), and where the gate underflows to exactly zero the kernel returns -0.0 with unbounded relative error. In-place operation (output == input) is supported on both legs; each element is read before it is written.

:::note
NaN propagates (TensorFlow Lite semantics): a NaN input element yields NaN at that output element at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast: propagation rides the final multiply x * gate  a NaN x makes the product NaN whatever the gate resolved to  rather than a compare-and-select that -ffinite-math-only could fold. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. +Inf returns +Inf (the gate is 1). -Inf returns NaN, not the mathematical limit 0: the gate is 0 there and (-Inf) * 0 is NaN by IEEE 754, the same result TFLite's float hard-swish reference produces; special-casing -Inf would put a per-element select in the hot loop for an input no finite model produces. The scalar and MVE legs agree on the NaN-ness and on +/-Inf; NaN payload bits may differ between legs.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input samples. |
| output | float32_t * | out | Pointer to the output samples. |
| size | int32_t | in | Number of elements to process. Must be at least 1. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:680`

## arm_nn_activation_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_nn_activation_f16(
    const float16_t *input,
    float16_t *output,
    int32_t size,
    arm_nn_activation_type_flt type,
    float16_t act_param
)
```

Elementwise activation.

:::note
The RELU, RELU6 and LEAKY_RELU legs classify NaN on the integer bit pattern (#380 / #382), so a NaN input comes back as NaN at every optimization level on the gated toolchains, including the shipped -Ofast. This holds on both the scalar and the MVE (cortex-m55) build paths; the MVE RELU/RELU6 legs restore the NaN lanes that vmaxnmq/vminnmq suppress. The MVE TANH leg returns a NaN input unchanged and keeps the sign of zero, also decided on the bit pattern (#635). The scalar TANH leg returns NaN where there is no hardware floating point (__ARM_FP undefined, e.g. Cortex-M0), where it too classifies NaN on the bit pattern (quieting a signalling NaN), and elsewhere only in builds without -ffinite-math-only. SIGMOID and HARDSWISH are outside this contract; see the per-helper notes in `Include/Internal/arm_nn_activation_flt.h`.

:::

:::note
The HARDSWISH leg's scalar helper (arm_nn_hardswish_scalar_f32, serving every build that does not take the MVE float path  no MVE float support, or MVE present but not used, e.g. under ARM_MATH_AUTOVECTORIZE) keeps the legacy separately rounded multiply-and-add gate and can differ by an ulp in the curved region from the standalone `arm_hard_swish_f32`, whose gate is a correctly rounded fma; the mux's MVE helper (arm_nn_vhardswish_mve_f32) uses vfmaq and agrees with that kernel. Callers that need bit-exact, leg-agreeing hard swish  or the documented NaN/Inf contract  should call `arm_hard_swish_f32` directly.

:::

:::note
The RELU, RELU6 and LEAKY_RELU legs classify NaN on the integer bit pattern (#380 / #382), so a NaN input comes back as NaN at every optimization level on the gated toolchains, including the shipped -Ofast. This holds uniformly across build paths: the scalar path serves every build without MVE float16 (and LEAKY_RELU on MVE builds too), while the MVE RELU/RELU6 legs (cortex-m55) restore the NaN lanes that vmaxnmq/vminnmq suppress, using the same integer-domain lane classification as the elementwise clamps. SIGMOID and HARDSWISH are outside this contract; see the per-helper notes in `Include/Internal/arm_nn_activation_flt.h`.

:::

:::note
TANH propagates NaN on both scalar and MVE paths, preserves the sign of zero, and maps +/-Inf to +/-1, including under -Ofast. NaN payload, sign and signaling state are not specified. The finite LUT interpolation may round differently across paths; bitwise scalar/MVE agreement is not required. Caller FP control settings are not changed.

:::

:::note
Both legs of the HARDSWISH mux evaluate natively in float16  the scalar helper (arm_nn_hardswish_scalar_f16) with a separately rounded multiply-and-add gate, the MVE helper (arm_nn_vhardswish_mve_f16) with a float16 vfmaq  so either can differ by an ulp from the scalar leg of the standalone `arm_hard_swish_f16`, which computes in float32 with an fma gate and rounds to float16 once. Callers that need the documented NaN/Inf contract should call `arm_hard_swish_f16` directly.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input samples. |
| output | float16_t * | out | Pointer to the output samples. |
| size | int32_t | in | Number of elements to process. |
| type | arm_nn_activation_type_flt | in | Activation selector. |
| act_param | float16_t | in | Extra activation parameter. Used for parameterized activations such as leaky ReLU. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2750`

## arm_prelu_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_prelu_f16(
    const cmsis_nn_dims *input_dims,
    const float16_t *input,
    const cmsis_nn_dims *alpha_dims,
    const float16_t *alpha,
    const cmsis_nn_dims *output_dims,
    float16_t *output
)
```

Parametric ReLU for float32 data.

Computes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Must equal output_dims. |
| input | const float16_t * | in | Pointer to the input tensor. |
| alpha_dims | const cmsis_nn_dims * | in | Alpha tensor dimensions. |
| alpha | const float16_t * | in | Pointer to the alpha (slope) tensor. |
| output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. |
| output | float16_t * | out | Pointer to the output tensor. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2759`

## arm_hard_swish_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_hard_swish_f16(const float16_t *input, float16_t *output, int32_t size)
```

Hard swish activation for float16 data.

Computes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise. The scalar leg widens each element to float32, evaluates the gate and the product there exactly as in `arm_hard_swish_f32`, and narrows only the final product, so it is single-rounded. The MVE (cortex-m55) leg evaluates the same expression in float16 throughout, scaling the gate by 1/6 before the product so that the multiplier stays in [0, 1]; it rounds the gate and the product separately and so can sit up to 2 float16 ulp away from the scalar leg in the curved region -3 < x < 3. The saturated regions are exact and identical on both legs (x >= 3 returns x bit-exactly, x <= -3 returns zero), as is the NaN/Inf behavior below; NaN lanes agree in NaN-ness but not necessarily in payload. In-place operation (output == input) is supported on both legs.

:::note
NaN and Inf behave as in `arm_hard_swish_f32`, at every optimization level on the gated toolchains (see docs/guides/verification.md): NaN propagates through the final multiply (NaN-ness only, not a particular payload), +Inf returns +Inf, and -Inf returns NaN because the gate is 0 there and (-Inf) * 0 is NaN by IEEE 754, matching TFLite's float hard-swish reference rather than the mathematical limit 0.

:::

:::note
Nothing in this kernel converts between half and single precision any more, and the float16 kernels that still do are not tied to a particular assembler: they go through `Include/Internal/arm_nn_vcvt_f16.h`, which emits the scalar form of VCVTB/VCVTT wherever the vector form would be mis-encoded (binutils below 2.43). Under CMake the probe measures the assembler in use and selects the form; a build that never runs it  the CMSIS-Pack `Source` Cvariant, `module.mk`, or a CMake project that wires its architecture flags where the probe cannot read them  falls back to the compiler major, which is right for every Arm GNU release and wrong only for a GCC 14 or newer driver paired by hand with an older binutils. Check `as --version` if you assembled that pair yourself. See docs/guides/toolchains.md.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input samples. |
| output | float16_t * | out | Pointer to the output samples. |
| size | int32_t | in | Number of elements to process. Must be at least 1. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:2803`
