# heliaCORE.LSTM

## arm_lstm_unidirectional_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_lstm_unidirectional_f32(
    const float32_t *input,
    float32_t *output,
    const cmsis_nn_lstm_params_f32 *params,
    cmsis_nn_lstm_context_f32 *buffers
)
```

Unidirectional LSTM inference.

:::note
On the MVE float path a NaN cell state yields a NaN hidden state, as tanh returns NaN unchanged there (#635). With cell clipping enabled the clip removes a NaN cell state first.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Pointer to the input sequence tensor. |
| output | float32_t * | out | Pointer to the output sequence tensor. |
| params | const cmsis_nn_lstm_params_f32 * | in | LSTM parameters and weights. |
| buffers | cmsis_nn_lstm_context_f32 * | in, out | Mutable LSTM scratch and state buffers. temp1 and temp2 are sized by `arm_lstm_unidirectional_f32_temp1_get_buffer_size()` / `arm_lstm_unidirectional_f32_temp2_get_buffer_size()`, which report 0: the float implementation never dereferences them and both may be NULL. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:1781`

## arm_gru_unidirectional_f32

`function` · `c`

```c
arm_cmsis_nn_status arm_gru_unidirectional_f32(
    const float32_t *input,
    float32_t *output,
    const cmsis_nn_gru_params_f32 *params,
    cmsis_nn_gru_context_f32 *buffers
)
```

Unidirectional GRU layer for float32 input, output and state.

Implements the reset-after GRU (Keras / TFLite default) when `params->reset_after` is non-zero, and the pre-reset variant otherwise. The hidden state is zero-initialised for the first time step, unless `buffers->hidden_state` is supplied for streaming state carry (`batch_size == 1`), in which case it seeds the initial state and receives the final hidden state on return.

:::note
NaN contract: a NaN in `input`, the previous hidden state, or the candidate gate's weight or bias reaches every output unit it feeds, on the scalar and MVE legs alike and at the shipped -Ofast: the MVE block re-establishes NaN after the table tanh with an integer-domain test that fast-math cannot elide (#251). A NaN confined to the update or reset gate's weight or bias does not reach the output: the scalar sigmoid maps NaN to 1.0 (see the note on arm_nn_sigmoid_scalar_f32 in `arm_nnsupportfunctions_flt.h`). NaN payloads and signs are not preserved on the MVE leg (default NaN, architectural). Inf follows the arithmetic.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float32_t * | in | Input sequence tensor. Must not overlap `output`: earlier outputs are re-read as the recurrent state for later time steps, so aliasing corrupts silently. |
| output | float32_t * | out | Output (hidden-state) sequence tensor. |
| params | const cmsis_nn_gru_params_f32 * | in | Struct describing the GRU operator. |
| buffers | cmsis_nn_gru_context_f32 * | in, out | Scratch buffers. May be NULL when `reset_after` != 0. temp1 is sized by `arm_gru_unidirectional_f32_temp1_get_buffer_size()`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, ARM_CMSIS_NN_ARG_ERROR otherwise. |

Source: `Include/arm_nnfunctions_flt.h:1812`

## arm_lstm_unidirectional_f32_temp1_get_buffer_size

`function` · `c`

```c
int32_t arm_lstm_unidirectional_f32_temp1_get_buffer_size(const cmsis_nn_lstm_params_f32 *lstm_params)
```

Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_f32()`.

:::note
This query reports its invalid input as -1, following the integer LSTM temp sizers (`arm_lstm_unidirectional_s8_temp1_get_buffer_size()`), not the 0 used by the float convolution and fully-connected queries in this header.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params_f32 * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f32` passed to `arm_lstm_unidirectional_f32()`. No field is read. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 0 for any non-NULL lstm_params, on every build target: the float32 implementation computes its gate values per hidden unit in automatics and never dereferences temp1 or temp2, so both context pointers may be NULL. Returns -1 only for a NULL lstm_params. The query exists so arena-sizing code can treat every LSTM variant alike; a future implementation that starts staging gate vectors would change this figure, so size from the query rather than hard-coding 0. |

Source: `Include/arm_nnfunctions_flt.h:1833`

## arm_lstm_unidirectional_f32_temp2_get_buffer_size

`function` · `c`

```c
int32_t arm_lstm_unidirectional_f32_temp2_get_buffer_size(const cmsis_nn_lstm_params_f32 *lstm_params)
```

Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_f32()`. The contract is identical to `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, and the answer is the same 0 (temp2 is likewise never dereferenced).

:::note
This query reports its invalid input as -1, following the integer LSTM temp sizers (`arm_lstm_unidirectional_s8_temp1_get_buffer_size()`), not the 0 used by the float convolution and fully-connected queries in this header.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params_f32 * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f32` passed to `arm_lstm_unidirectional_f32()`. No field is read. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 0 for any non-NULL lstm_params, on every build target: the float32 implementation computes its gate values per hidden unit in automatics and never dereferences temp1 or temp2, so both context pointers may be NULL. Returns -1 only for a NULL lstm_params. The query exists so arena-sizing code can treat every LSTM variant alike; a future implementation that starts staging gate vectors would change this figure, so size from the query rather than hard-coding 0. |

Source: `Include/arm_nnfunctions_flt.h:1842`

## arm_gru_unidirectional_f32_temp1_get_buffer_size

`function` · `c`

```c
int32_t arm_gru_unidirectional_f32_temp1_get_buffer_size(const cmsis_nn_gru_params_f32 *gru_params)
```

Get size of the temp1 scratch buffer required by `arm_gru_unidirectional_f32()`.

:::note
On the pre-reset path a 0 is only returned for the degenerate hidden_size == 0, which `arm_gru_unidirectional_f32()` rejects with ARM_CMSIS_NN_ARG_ERROR before any buffer access - so a 0 there never corresponds to a runnable call.

:::

:::note
This query reports an out-of-range shape as -1, following the integer LSTM temp sizers, not the 0 used by the float convolution and fully-connected queries in this header.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| gru_params | const cmsis_nn_gru_params_f32 * | in | GRU operator parameters, i.e. the same `cmsis_nn_gru_params_f32` passed to `arm_gru_unidirectional_f32()`. Only reset_after and hidden_size are read. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes: hidden_size * sizeof(float32_t) when reset_after == 0 (the pre-reset formulation stages the r . h_prev vector in temp1; the vector is reused across batches and time steps, so neither batch_size nor time_steps enters), and 0 when reset_after != 0 (temp1 is never dereferenced and may be NULL). Returns -1 if gru_params is NULL, if hidden_size is negative, or if the byte count would not fit in an int32_t. The figure and the range checks are the same on every build target. |

Source: `Include/arm_nnfunctions_flt.h:1863`

## arm_lstm_unidirectional_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_lstm_unidirectional_f16(
    const float16_t *input,
    float16_t *output,
    const cmsis_nn_lstm_params_f16 *params,
    cmsis_nn_lstm_context_f16 *buffers
)
```

Unidirectional LSTM inference.

:::note
On the MVE float path a NaN cell state yields a NaN hidden state, as tanh returns NaN unchanged there (#635). With cell clipping enabled the clip removes a NaN cell state first.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Pointer to the input sequence tensor. |
| output | float16_t * | out | Pointer to the output sequence tensor. |
| params | const cmsis_nn_lstm_params_f16 * | in | LSTM parameters and weights. |
| buffers | cmsis_nn_lstm_context_f16 * | in, out | Mutable LSTM scratch and state buffers. temp1 and temp2 are sized by `arm_lstm_unidirectional_f32_temp1_get_buffer_size()` / `arm_lstm_unidirectional_f32_temp2_get_buffer_size()`, which report 0: the float implementation never dereferences them and both may be NULL. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | `ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments. |

Source: `Include/arm_nnfunctions_flt.h:3558`

## arm_gru_unidirectional_f16

`function` · `c`

```c
arm_cmsis_nn_status arm_gru_unidirectional_f16(
    const float16_t *input,
    float16_t *output,
    const cmsis_nn_gru_params_f16 *params,
    cmsis_nn_gru_context_f16 *buffers
)
```

Unidirectional GRU layer for float16 input, output and state.

Implements the reset-after GRU (Keras / TFLite default) when `params->reset_after` is non-zero, and the pre-reset variant otherwise. The hidden state is zero-initialised for the first time step, unless `buffers->hidden_state` is supplied for streaming state carry (`batch_size == 1`), in which case it seeds the initial state and receives the final hidden state on return.

:::note
NaN contract: a NaN in `input`, the previous hidden state, or the candidate gate's weight or bias reaches every output unit it feeds, on the scalar and MVE legs alike and at the shipped -Ofast: the MVE block re-establishes NaN after the table tanh with an integer-domain test that fast-math cannot elide (#251). A NaN confined to the update or reset gate's weight or bias does not reach the output: the scalar sigmoid maps NaN to 1.0 (see the note on arm_nn_sigmoid_scalar_f32 in `arm_nnsupportfunctions_flt.h`). NaN payloads and signs are not preserved on the MVE leg (default NaN, architectural). Inf follows the arithmetic.

:::

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| input | const float16_t * | in | Input sequence tensor. Must not overlap `output`: earlier outputs are re-read as the recurrent state for later time steps, so aliasing corrupts silently. |
| output | float16_t * | out | Output (hidden-state) sequence tensor. |
| params | const cmsis_nn_gru_params_f16 * | in | Struct describing the GRU operator. |
| buffers | cmsis_nn_gru_context_f16 * | in, out | Scratch buffers. May be NULL when `reset_after` != 0. temp1 is sized by `arm_gru_unidirectional_f16_temp1_get_buffer_size()`. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | ARM_CMSIS_NN_SUCCESS on success, ARM_CMSIS_NN_ARG_ERROR otherwise. |

Source: `Include/arm_nnfunctions_flt.h:3589`

## arm_lstm_unidirectional_f16_temp1_get_buffer_size

`function` · `c`

```c
int32_t arm_lstm_unidirectional_f16_temp1_get_buffer_size(const cmsis_nn_lstm_params_f16 *lstm_params)
```

Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_f16()`. The contract is identical to `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, and the answer is the same 0 on every build target (the float16 implementation likewise never dereferences temp1 or temp2, which may both be NULL).

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params_f16 * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f16` passed to `arm_lstm_unidirectional_f16()`. No field is read. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 0 for any non-NULL lstm_params, -1 for a NULL lstm_params. |

Source: `Include/arm_nnfunctions_flt.h:3605`

## arm_lstm_unidirectional_f16_temp2_get_buffer_size

`function` · `c`

```c
int32_t arm_lstm_unidirectional_f16_temp2_get_buffer_size(const cmsis_nn_lstm_params_f16 *lstm_params)
```

Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_f16()`. The contract is identical to `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, and the answer is the same 0.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| lstm_params | const cmsis_nn_lstm_params_f16 * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f16` passed to `arm_lstm_unidirectional_f16()`. No field is read. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | 0 for any non-NULL lstm_params, -1 for a NULL lstm_params. |

Source: `Include/arm_nnfunctions_flt.h:3613`

## arm_gru_unidirectional_f16_temp1_get_buffer_size

`function` · `c`

```c
int32_t arm_gru_unidirectional_f16_temp1_get_buffer_size(const cmsis_nn_gru_params_f16 *gru_params)
```

Get size of the temp1 scratch buffer required by `arm_gru_unidirectional_f16()`. See `arm_gru_unidirectional_f32_temp1_get_buffer_size()` for the -1-on-invalid contract and the pre-reset degenerate-0 note; both apply here unchanged.

**Parameters**

| Name | Type | Direction | Description |
| --- | --- | --- | --- |
| gru_params | const cmsis_nn_gru_params_f16 * | in | GRU operator parameters, i.e. the same `cmsis_nn_gru_params_f16` passed to `arm_gru_unidirectional_f16()`. Only reset_after and hidden_size are read. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  |  | Required buffer size in bytes: hidden_size * sizeof(float16_t) when reset_after == 0, 0 when reset_after != 0 (temp1 is never dereferenced and may be NULL). Half the figure `arm_gru_unidirectional_f32_temp1_get_buffer_size()` returns for the same shape - sizing an f16 layer with the f32 query over-allocates, and the reverse under-allocates. |

Source: `Include/arm_nnfunctions_flt.h:3628`
