Skip to content
heliaCORE
API reference
HELIA HUB

Internal support

Internal Support functions. Not intended to be called direclty by a CMSIS-NN user.

Machine-readable model

attribute

LUT for 2^(i/256) used by the float32 LUT softmax approximation.

Include/arm_nnsupportfunctions_flt.h:78

const float32_t arm_nn_exp2_lut_f32[257]

LUT for 2^(i/256) used by the float32 LUT softmax approximation.

Stores 257 samples for i = 0..256 so interpolation can safely read lut[idx + 1] while indexing the 256 fractional segments.

function

Floor of x as an int32t.

Include/arm_nnsupportfunctions_flt.h:92

static int32_t arm_nn_softmax_floor_to_int_f32(float32_t x)

Floor of x as an int32_t.

Precondition: x must already be reduced to the int32_t range and must not be NaN the float-to-int conversion below is undefined otherwise. The only caller, arm_nn_softmax_exp_lut_f32(), guarantees this by clamping its input to [-80, 80] (NaN included, see there) before scaling by log2(e), which bounds x to +/-116.

Parameters of arm_nn_softmax_floor_to_int_f32
NameTypeDirectionDescription
xfloat32_tinValue to floor.
Returns of arm_nn_softmax_floor_to_int_f32
Description
Largest int32_t not greater than `x`.
function

Reinterpret a 32-bit pattern as a float32.

Include/arm_nnsupportfunctions_flt.h:104

static float32_t arm_nn_softmax_fp32_from_bits(uint32_t bits)

Reinterpret a 32-bit pattern as a float32.

Parameters of arm_nn_softmax_fp32_from_bits
NameTypeDirectionDescription
bitsuint32_tinIEEE-754 binary32 bit pattern.
Returns of arm_nn_softmax_fp32_from_bits
Description
The float32 value with the bit pattern `bits`.
function

Compute 2^n as a float32 by building the exponent field directly.

Include/arm_nnsupportfunctions_flt.h:121

static float32_t arm_nn_softmax_exp2i_f32(int32_t n)

Compute 2^n as a float32 by building the exponent field directly.

Parameters of arm_nn_softmax_exp2i_f32
NameTypeDirectionDescription
nint32_tinInteger exponent. Clamped to the normal float32 exponent range `[-126, 127]`.
Returns of arm_nn_softmax_exp2i_f32
Description
`2^n` as a float32.
function

Taylor/Estrin exp approximation for float32 softmax helpers.

Include/arm_nnsupportfunctions_flt.h:147

static float32_t arm_nn_softmax_exp_taylor_f32(float32_t x)

Taylor/Estrin exp approximation for float32 softmax helpers.

The polynomial is evaluated on r in [-ln2/2, ln2/2]. Coefficients come from the Maclaurin series of exp(r): exp(r) ~= 1 + r + r^2/2! + r^3/3! + r^4/4! + r^5/5! + r^6/6! Grouped via Estrin to reduce dependency depth: p = (1 + r) + r^2*(1/2 + r/6) + r^4*(1/24 + r/120) + r^6*(1/720)

Range reduction follows: x = n * ln(2) + r, exp(x) = exp(r) * 2^n

Parameters of arm_nn_softmax_exp_taylor_f32
NameTypeDirectionDescription
xfloat32_tinExponent argument. Clamped to `[-80, 80]` before evaluation.
Returns of arm_nn_softmax_exp_taylor_f32
Description
Approximation of `exp(x)`.
function

LUT-based exp approximation for float32 softmax helpers.

Include/arm_nnsupportfunctions_flt.h:184

static float32_t arm_nn_softmax_exp_lut_f32(float32_t x)

LUT-based exp approximation for float32 softmax helpers.

Splits x * log2(e) into an integer part handled by arm_nn_softmax_exp2i_f32() and a fractional part interpolated linearly from arm_nn_exp2_lut_f32.

Parameters of arm_nn_softmax_exp_lut_f32
NameTypeDirectionDescription
xfloat32_tinExponent argument. Clamped to `[-80, 80]` before evaluation; NaN is flushed to `80`.
Returns of arm_nn_softmax_exp_lut_f32
Description
Approximation of `exp(x)`.
function

Scalar exp approximation used by the float32 softmax paths.

Include/arm_nnsupportfunctions_flt.h:254

static float32_t arm_nn_softmax_exp_scalar_f32(float32_t x)

Scalar exp approximation used by the float32 softmax paths.

Dispatches to arm_nn_softmax_exp_taylor_f32() when ARM_NN_USE_EXP_TAYLOR is defined and to arm_nn_softmax_exp_lut_f32() otherwise.

Parameters of arm_nn_softmax_exp_scalar_f32
NameTypeDirectionDescription
xfloat32_tinExponent argument.
Returns of arm_nn_softmax_exp_scalar_f32
Description
Approximation of `exp(x)`.
attribute

LUT for tanh(x) sampled over x in [0, 6] for float32 helpers.

Include/arm_nnsupportfunctions_flt.h:276

const float32_t arm_nn_tanh_lut_f32[385]

LUT for tanh(x) sampled over x in [0, 6] for float32 helpers.

Stores 385 samples so interpolation can safely read lut[idx + 1] while indexing the 384 fractional segments across the interval. The grid spacing (6/384 == 1/64) matches the earlier 257-entry [0, 4] table, so entries 0..256 are bit-identical to it and the index multiplier is unchanged. Generated by scripts/gen_tanh_lut_f32.py.

function

Copy a float32 vector.

Include/arm_nnsupportfunctions_flt.h:311

static void arm_memcpy_f32(float32_t *dst, const float32_t *src, uint32_t block_size)

Copy a float32 vector.

Parameters of arm_memcpy_f32
NameTypeDirectionDescription
dstfloat32_t *outDestination buffer.
srcconst float32_t *inSource buffer.
block_sizeuint32_tinNumber of elements to copy.
function

Set a float32 vector to a constant value.

Include/arm_nnsupportfunctions_flt.h:334

static void arm_memset_f32(float32_t *dst, const float32_t val, uint32_t block_size)

Set a float32 vector to a constant value.

Parameters of arm_memset_f32
NameTypeDirectionDescription
dstfloat32_t *outDestination buffer.
valconst float32_tinFill value.
block_sizeuint32_tinNumber of elements to write.
function

Specialized NHWC depthwise 1D kernel for k=3, chmult=1 (float32).

Include/arm_nnsupportfunctions_flt.h:367

void arm_nn_depthwise_conv1d_k3_nhwc_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel,
const float32_t *b,
float32_t *out,
int32_t out_w
)

Specialized NHWC depthwise 1D kernel for k=3, ch_mult=1 (float32).

Parameters of arm_nn_depthwise_conv1d_k3_nhwc_f32
NameTypeDirectionDescription
x_nhwcconst float32_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input (and output) channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernelconst float32_t *inDepthwise weights with shape `[3][in_c]`.
bconst float32_t *inOptional bias vector of `in_c` elements. May be NULL.
outfloat32_t *outOutput row in NHWC layout with shape `[out_w][in_c]`.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+2`.
function

Specialized NHWC 1D convolution kernel for k=5 (float32).

Include/arm_nnsupportfunctions_flt.h:387

void arm_nn_conv1d_k5_nhwc_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel,
const float32_t *b,
float32_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=5 (float32).

Parameters of arm_nn_conv1d_k5_nhwc_f32
NameTypeDirectionDescription
x_nhwcconst float32_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernelconst float32_t *inWeights with shape `[out_c][5][in_c]`.
bconst float32_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat32_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+4`.
function

Specialized NHWC 1D convolution kernel for k=5 (float32, packed weights).

Include/arm_nnsupportfunctions_flt.h:411

void arm_nn_conv1d_k5_packed_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel_packed,
const float32_t *b,
float32_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=5 (float32, packed weights).

The packed kernel uses the same NTxN RHS layout as arm_nn_mat_mult_nt_n_packed_f32, i.e. [(5 * in_c)][out_c_block_of_4].

Parameters of arm_nn_conv1d_k5_packed_f32
NameTypeDirectionDescription
x_nhwcconst float32_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernel_packedconst float32_t *inWeights packed in output-channel blocks of 4 as described above.
bconst float32_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat32_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+4`.
function

Specialized NHWC 1D convolution kernel for k=3 (float32).

Include/arm_nnsupportfunctions_flt.h:432

void arm_nn_conv1d_k3_nhwc_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel,
const float32_t *b,
float32_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=3 (float32).

Parameters of arm_nn_conv1d_k3_nhwc_f32
NameTypeDirectionDescription
x_nhwcconst float32_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernelconst float32_t *inWeights with shape `[out_c][3][in_c]`.
bconst float32_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat32_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+2`.
function

Specialized NHWC 1D convolution kernel for k=3 (float32, packed weights).

Include/arm_nnsupportfunctions_flt.h:456

void arm_nn_conv1d_k3_packed_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float32_t *kernel_packed,
const float32_t *b,
float32_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=3 (float32, packed weights).

The packed kernel uses the same NTxN RHS layout as arm_nn_mat_mult_nt_n_packed_f32, i.e. [(3 * in_c)][out_c_block_of_4].

Parameters of arm_nn_conv1d_k3_packed_f32
NameTypeDirectionDescription
x_nhwcconst float32_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernel_packedconst float32_t *inWeights packed in output-channel blocks of 4 as described above.
bconst float32_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat32_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+2`.
function

Specialized NHWC max-pool 1D kernel for k=3, s=3 (float32).

Include/arm_nnsupportfunctions_flt.h:474

void arm_nn_maxpool1d_k3s3_nhwc_f32(const float32_t *x_nhwc, int32_t in_c, int32_t in_w, float32_t *out, int32_t out_w)

Specialized NHWC max-pool 1D kernel for k=3, s=3 (float32).

Parameters of arm_nn_maxpool1d_k3s3_nhwc_f32
NameTypeDirectionDescription
x_nhwcconst float32_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of channels.
in_wint32_tinInput width. Currently unused by the kernel.
outfloat32_t *outOutput row in NHWC layout with shape `[out_w][in_c]`.
out_wint32_tinOutput width. Output position `ow` reads input positions `3*ow..3*ow+2`.
function

Specialized NHWC max-pool 1D kernel for k=2, s=2 without output clamp (float32).

Include/arm_nnsupportfunctions_flt.h:489

void arm_nn_maxpool1d_k2s2_nhwc_noclip_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
float32_t *out,
int32_t out_w
)

Specialized NHWC max-pool 1D kernel for k=2, s=2 without output clamp (float32).

Parameters of arm_nn_maxpool1d_k2s2_nhwc_noclip_f32
NameTypeDirectionDescription
x_nhwcconst float32_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of channels.
in_wint32_tinInput width. Currently unused by the kernel.
outfloat32_t *outOutput row in NHWC layout with shape `[out_w][in_c]`.
out_wint32_tinOutput width. Output position `ow` reads input positions `2*ow..2*ow+1`.
function

Specialized NHWC max-pool 1D kernel for k=2, s=2 with clamp (float32).

Include/arm_nnsupportfunctions_flt.h:506

void arm_nn_maxpool1d_k2s2_nhwc_f32(
const float32_t *x_nhwc,
int32_t in_c,
int32_t in_w,
float32_t *out,
int32_t out_w,
float32_t act_min,
float32_t act_max
)

Specialized NHWC max-pool 1D kernel for k=2, s=2 with clamp (float32).

Parameters of arm_nn_maxpool1d_k2s2_nhwc_f32
NameTypeDirectionDescription
x_nhwcconst float32_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of channels.
in_wint32_tinInput width. Currently unused by the kernel.
outfloat32_t *outOutput row in NHWC layout with shape `[out_w][in_c]`.
out_wint32_tinOutput width. Output position `ow` reads input positions `2*ow..2*ow+1`.
act_minfloat32_tinLower clamp bound applied to `out`.
act_maxfloat32_tinUpper clamp bound applied to `out`.
function

Matrix multiply with non-transposed lhs and transposed rhs rows (float32).

Include/arm_nnsupportfunctions_flt.h:529

arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f32(
const float32_t *lhs,
const float32_t *rhs,
const float32_t *bias,
float32_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float32_t activation_min,
float32_t activation_max
)

Matrix multiply with non-transposed lhs and transposed rhs rows (float32).

Parameters of arm_nn_mat_mult_nt_t_f32
NameTypeDirectionDescription
lhsconst float32_t *inLeft-hand matrix stored row-major.
rhsconst float32_t *inRight-hand matrix stored row-major, one row per output channel.
biasconst float32_t *inOptional bias vector.
dstfloat32_t *outOutput matrix.
lhs_rowsint32_tinNumber of rows in `lhs`.
rhs_rowsint32_tinNumber of rows in `rhs`.
rhs_colsint32_tinNumber of columns in `rhs`.
row_address_offsetint32_tinOutput row stride, expressed in elements.
activation_minfloat32_tinLower clamp bound.
activation_maxfloat32_tinUpper clamp bound.
Returns of arm_nn_mat_mult_nt_t_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Matrix multiply with non-transposed lhs and packed non-transposed rhs (float32).

Include/arm_nnsupportfunctions_flt.h:556

arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f32(
const float32_t *lhs,
const float32_t *rhs_packed,
const float32_t *bias,
float32_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float32_t activation_min,
float32_t activation_max
)

Matrix multiply with non-transposed lhs and packed non-transposed rhs (float32).

Parameters of arm_nn_mat_mult_nt_n_packed_f32
NameTypeDirectionDescription
lhsconst float32_t *inLeft-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`.
rhs_packedconst float32_t *inRight-hand matrix with logical shape `[rhs_cols, rhs_rows]`, packed in column blocks of 4. The final block uses the same packed stride and inactive tail lanes are ignored.
biasconst float32_t *inOptional bias vector.
dstfloat32_t *outOutput matrix.
lhs_rowsint32_tinNumber of rows in `lhs`.
rhs_rowsint32_tinNumber of logical output columns in the unpacked rhs matrix.
rhs_colsint32_tinShared reduction dimension `K`.
row_address_offsetint32_tinOutput row stride, expressed in elements.
activation_minfloat32_tinLower clamp bound.
activation_maxfloat32_tinUpper clamp bound.
Returns of arm_nn_mat_mult_nt_n_packed_f32
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Pack a single convolution patch into one row of a contiguous float32 patch matrix.

Include/arm_nnsupportfunctions_flt.h:590

void arm_nn_pack_conv_patch_f32(
const float32_t *input,
int32_t in_h,
int32_t in_w,
int32_t in_c,
int32_t kernel_h,
int32_t kernel_w,
int32_t stride_h,
int32_t stride_w,
int32_t pad_h,
int32_t pad_w,
int32_t dilation_h,
int32_t dilation_w,
int32_t out_y,
int32_t out_x,
float32_t pad_value,
float32_t *patch_row
)

Pack a single convolution patch into one row of a contiguous float32 patch matrix.

Developers familiar with im2row/im2col terminology can think of this as packing one output patch into one row.

Parameters of arm_nn_pack_conv_patch_f32
NameTypeDirectionDescription
inputconst float32_t *inInput tensor for one batch in NHWC layout with shape `[in_h][in_w][in_c]`.
in_hint32_tinInput height.
in_wint32_tinInput width.
in_cint32_tinNumber of input channels.
kernel_hint32_tinKernel height.
kernel_wint32_tinKernel width.
stride_hint32_tinVertical stride.
stride_wint32_tinHorizontal stride.
pad_hint32_tinTop padding.
pad_wint32_tinLeft padding.
dilation_hint32_tinVertical dilation.
dilation_wint32_tinHorizontal dilation.
out_yint32_tinOutput row index of the patch to pack.
out_xint32_tinOutput column index of the patch to pack.
pad_valuefloat32_tinValue written for taps that fall outside the input.
patch_rowfloat32_t *outDestination row of `kernel_h * kernel_w * in_c` elements, ordered `[kernel_h][kernel_w][in_c]`.
function

Specialized softmax helper for a single float32 row of length 2.

Include/arm_nnsupportfunctions_flt.h:613

void arm_nn_softmax_1x2_f32(const float32_t *in, float32_t *out)

Specialized softmax helper for a single float32 row of length 2.

Parameters of arm_nn_softmax_1x2_f32
NameTypeDirectionDescription
inconst float32_t *inPointer to two contiguous float32 input values.
outfloat32_t *outPointer to two contiguous float32 output values.
macro

Blockwise float16 accumulation on the MVE legs (AmbiqAI/ns-cmsis-nn#586).

Include/arm_nnsupportfunctions_flt.h:626

#define ARM_NN_F16_ACC_BLOCK (32)

Blockwise float16 accumulation on the MVE legs (AmbiqAI/ns-cmsis-nn#586).

A float16 accumulator lane sums at most ARM_NN_F16_ACC_BLOCK taps, in the kernel’s tap order, before its partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once. The _acc16 entries instantiate the same kernel bodies with ARM_NN_F16_ACC_BLOCK_NONE, which never folds.

attribute

Polynomial coefficients used by the float16 MVE exp approximation.

Include/arm_nnsupportfunctions_flt.h:636

const float32_t arm_nn_exp_poly_coeffs_f16[8]

Polynomial coefficients used by the float16 MVE exp approximation.

The float16 MVE helper evaluates the polynomial in widened float32 lanes, but it uses a dedicated coefficient table to keep the float16 path isolated from the float32 feature gate and softmax support stack.

attribute

Quantized binary16 LUT for 2^(i/256) used by float16 helpers.

Include/arm_nnsupportfunctions_flt.h:644

const uint16_t arm_nn_exp2_lut_f16[257]

Quantized binary16 LUT for 2^(i/256) used by float16 helpers.

Stores 257 samples for i = 0..256 so interpolation can safely read lut[idx + 1] while indexing the 256 fractional segments.

attribute

Quantized binary16 LUT for tanh(x) with x in [0, 4].

Include/arm_nnsupportfunctions_flt.h:652

const uint16_t arm_nn_tanh_lut_f16[257]

Quantized binary16 LUT for tanh(x) with x in [0, 4].

Stores 257 samples so interpolation can safely read lut[idx + 1] while indexing the 256 fractional segments across the interval.

function

Reinterpret a 16-bit pattern as a float16.

Include/arm_nnsupportfunctions_flt.h:660

static float16_t arm_nn_softmax_fp16_from_bits(uint16_t bits)

Reinterpret a 16-bit pattern as a float16.

Parameters of arm_nn_softmax_fp16_from_bits
NameTypeDirectionDescription
bitsuint16_tinIEEE-754 binary16 bit pattern.
Returns of arm_nn_softmax_fp16_from_bits
Description
The float16 value with the bit pattern `bits`.
function

Floor of x as an int32t.

Include/arm_nnsupportfunctions_flt.h:677

static int32_t arm_nn_softmax_floor_to_int_f16(float16_t x)

Floor of x as an int32_t.

Parameters of arm_nn_softmax_floor_to_int_f16
NameTypeDirectionDescription
xfloat16_tinValue to floor. Must be finite and within the int32_t range.
Returns of arm_nn_softmax_floor_to_int_f16
Description
Largest int32_t not greater than `x`.
function

Compute 2^n as a float16 by building the exponent field directly.

Include/arm_nnsupportfunctions_flt.h:690

static float16_t arm_nn_softmax_exp2i_f16(int32_t n)

Compute 2^n as a float16 by building the exponent field directly.

Parameters of arm_nn_softmax_exp2i_f16
NameTypeDirectionDescription
nint32_tinInteger exponent. Clamped to the normal float16 exponent range `[-14, 15]`.
Returns of arm_nn_softmax_exp2i_f16
Description
`2^n` as a float16.
function

Taylor/Estrin exp approximation for float16 softmax helpers.

Include/arm_nnsupportfunctions_flt.h:710

static float16_t arm_nn_softmax_exp_taylor_f16(float16_t x)

Taylor/Estrin exp approximation for float16 softmax helpers.

The evaluation uses float32 intermediates to keep the approximation stable, but it is fully independent from the float32 softmax support tables.

Parameters of arm_nn_softmax_exp_taylor_f16
NameTypeDirectionDescription
xfloat16_tinExponent argument. Clamped to `[-80, 80]` before evaluation.
Returns of arm_nn_softmax_exp_taylor_f16
Description
Approximation of `exp(x)`.
function

LUT-based exp approximation for float16 softmax helpers.

Include/arm_nnsupportfunctions_flt.h:746

static float16_t arm_nn_softmax_exp_lut_f16(float16_t x)

LUT-based exp approximation for float16 softmax helpers.

Splits x * log2(e) into an integer part handled by arm_nn_softmax_exp2i_f16() and a fractional part interpolated linearly from arm_nn_exp2_lut_f16, using float32 intermediates.

Parameters of arm_nn_softmax_exp_lut_f16
NameTypeDirectionDescription
xfloat16_tinExponent argument. Clamped to `[-80, 80]` before evaluation.
Returns of arm_nn_softmax_exp_lut_f16
Description
Approximation of `exp(x)`.
function

Scalar exp approximation used by the float16 softmax paths.

Include/arm_nnsupportfunctions_flt.h:787

static float16_t arm_nn_softmax_exp_scalar_f16(float16_t x)

Scalar exp approximation used by the float16 softmax paths.

Dispatches to arm_nn_softmax_exp_taylor_f16() when ARM_NN_USE_EXP_TAYLOR is defined and to arm_nn_softmax_exp_lut_f16() otherwise.

Parameters of arm_nn_softmax_exp_scalar_f16
NameTypeDirectionDescription
xfloat16_tinExponent argument.
Returns of arm_nn_softmax_exp_scalar_f16
Description
Approximation of `exp(x)`.
function

Copy a float16 vector.

Include/arm_nnsupportfunctions_flt.h:979

static void arm_memcpy_f16(float16_t *dst, const float16_t *src, uint32_t block_size)

Copy a float16 vector.

Parameters of arm_memcpy_f16
NameTypeDirectionDescription
dstfloat16_t *outDestination buffer.
srcconst float16_t *inSource buffer.
block_sizeuint32_tinNumber of elements to copy.
function

Set a float16 vector to a constant value.

Include/arm_nnsupportfunctions_flt.h:1002

static void arm_memset_f16(float16_t *dst, const float16_t val, uint32_t block_size)

Set a float16 vector to a constant value.

Parameters of arm_memset_f16
NameTypeDirectionDescription
dstfloat16_t *outDestination buffer.
valconst float16_tinFill value.
block_sizeuint32_tinNumber of elements to write.
function

Specialized NHWC depthwise 2x5 kernel (float16).

Include/arm_nnsupportfunctions_flt.h:1039

void arm_nn_depthwise_conv2x5_nhwc_f16(
const float16_t *x_nhwc,
int32_t batches,
int32_t in_c,
int32_t in_w,
int32_t ch_mult,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_w,
float16_t act_min,
float16_t act_max
)

Specialized NHWC depthwise 2x5 kernel (float16).

Parameters of arm_nn_depthwise_conv2x5_nhwc_f16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput tensor in NHWC layout with shape `[batches][2][in_w][in_c]`.
batchesint32_tinNumber of batches.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width.
ch_multint32_tinChannel multiplier; the output has `in_c * ch_mult` channels.
kernelconst float16_t *inDepthwise weights with shape `[2][5][in_c * ch_mult]`.
bconst float16_t *inOptional bias vector of `in_c * ch_mult` elements. May be NULL.
outfloat16_t *outOutput tensor in NHWC layout with shape `[batches][1][out_w][in_c * ch_mult]`.
out_wint32_tinOutput width. Output position `ow` reads input columns `ow..ow+4`.
act_minfloat16_tinLower clamp bound applied to `out`.
act_maxfloat16_tinUpper clamp bound applied to `out`.
function

Specialized NHWC depthwise 1D kernel for k=3, chmult=1 (float32).

Include/arm_nnsupportfunctions_flt.h:1054

void arm_nn_depthwise_conv1d_k3_nhwc_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_w
)

Specialized NHWC depthwise 1D kernel for k=3, ch_mult=1 (float32).

Parameters of arm_nn_depthwise_conv1d_k3_nhwc_f16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input (and output) channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernelconst float16_t *inDepthwise weights with shape `[3][in_c]`.
bconst float16_t *inOptional bias vector of `in_c` elements. May be NULL.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][in_c]`.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+2`.
function

Specialized NHWC 1D convolution kernel for k=5 (float32).

Include/arm_nnsupportfunctions_flt.h:1073

void arm_nn_conv1d_k5_nhwc_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=5 (float32).

Parameters of arm_nn_conv1d_k5_nhwc_f16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernelconst float16_t *inWeights with shape `[out_c][5][in_c]`.
bconst float16_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+4`.
function

Specialized NHWC 1D convolution kernel for k=5 (float32).

Include/arm_nnsupportfunctions_flt.h:1088

void arm_nn_conv1d_k5_nhwc_f16_acc16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=5 (float32).

Parameters of arm_nn_conv1d_k5_nhwc_f16_acc16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernelconst float16_t *inWeights with shape `[out_c][5][in_c]`.
bconst float16_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+4`.
function

Specialized NHWC 1D convolution kernel for k=5 (float16, packed weights).

Include/arm_nnsupportfunctions_flt.h:1118

void arm_nn_conv1d_k5_packed_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel_packed,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=5 (float16, packed weights).

The packed kernel uses the same NTxN RHS layout as arm_nn_mat_mult_nt_n_packed_f16, i.e. [(5 * in_c)][out_c_block_of_8].

Parameters of arm_nn_conv1d_k5_packed_f16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernel_packedconst float16_t *inWeights packed in output-channel blocks of 8 as described above.
bconst float16_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+4`.
function

Specialized NHWC 1D convolution kernel for k=5 (float16, packed weights).

Include/arm_nnsupportfunctions_flt.h:1133

void arm_nn_conv1d_k5_packed_f16_acc16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel_packed,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=5 (float16, packed weights).

The packed kernel uses the same NTxN RHS layout as arm_nn_mat_mult_nt_n_packed_f16, i.e. [(5 * in_c)][out_c_block_of_8].

Parameters of arm_nn_conv1d_k5_packed_f16_acc16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernel_packedconst float16_t *inWeights packed in output-channel blocks of 8 as described above.
bconst float16_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+4`.
function

Specialized NHWC 1D convolution kernel for k=3 (float32).

Include/arm_nnsupportfunctions_flt.h:1153

void arm_nn_conv1d_k3_nhwc_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=3 (float32).

Parameters of arm_nn_conv1d_k3_nhwc_f16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernelconst float16_t *inWeights with shape `[out_c][3][in_c]`.
bconst float16_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+2`.
function

Specialized NHWC 1D convolution kernel for k=3 (float32).

Include/arm_nnsupportfunctions_flt.h:1168

void arm_nn_conv1d_k3_nhwc_f16_acc16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=3 (float32).

Parameters of arm_nn_conv1d_k3_nhwc_f16_acc16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernelconst float16_t *inWeights with shape `[out_c][3][in_c]`.
bconst float16_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+2`.
function

Specialized NHWC 1D convolution kernel for k=3 (float16, packed weights).

Include/arm_nnsupportfunctions_flt.h:1198

void arm_nn_conv1d_k3_packed_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel_packed,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=3 (float16, packed weights).

The packed kernel uses the same NTxN RHS layout as arm_nn_mat_mult_nt_n_packed_f16, i.e. [(3 * in_c)][out_c_block_of_8].

Parameters of arm_nn_conv1d_k3_packed_f16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernel_packedconst float16_t *inWeights packed in output-channel blocks of 8 as described above.
bconst float16_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+2`.
function

Specialized NHWC 1D convolution kernel for k=3 (float16, packed weights).

Include/arm_nnsupportfunctions_flt.h:1213

void arm_nn_conv1d_k3_packed_f16_acc16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
const float16_t *kernel_packed,
const float16_t *b,
float16_t *out,
int32_t out_c,
int32_t out_w
)

Specialized NHWC 1D convolution kernel for k=3 (float16, packed weights).

The packed kernel uses the same NTxN RHS layout as arm_nn_mat_mult_nt_n_packed_f16, i.e. [(3 * in_c)][out_c_block_of_8].

Parameters of arm_nn_conv1d_k3_packed_f16_acc16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of input channels.
in_wint32_tinInput width. Currently unused by the kernel.
kernel_packedconst float16_t *inWeights packed in output-channel blocks of 8 as described above.
bconst float16_t *inOptional bias vector of `out_c` elements. May be NULL.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][out_c]`.
out_cint32_tinNumber of output channels.
out_wint32_tinOutput width. Output position `ow` reads input positions `ow..ow+2`.
function

Specialized NHWC max-pool 1D kernel for k=3, s=3 (float16).

Include/arm_nnsupportfunctions_flt.h:1227

void arm_nn_maxpool1d_k3s3_nhwc_f16(const float16_t *x_nhwc, int32_t in_c, int32_t in_w, float16_t *out, int32_t out_w)

Specialized NHWC max-pool 1D kernel for k=3, s=3 (float16).

Parameters of arm_nn_maxpool1d_k3s3_nhwc_f16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of channels.
in_wint32_tinInput width. Currently unused by the kernel.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][in_c]`.
out_wint32_tinOutput width. Output position `ow` reads input positions `3*ow..3*ow+2`.
function

Specialized NHWC max-pool 1D kernel for k=2, s=2 without output clamp (float16).

Include/arm_nnsupportfunctions_flt.h:1238

void arm_nn_maxpool1d_k2s2_nhwc_noclip_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
float16_t *out,
int32_t out_w
)

Specialized NHWC max-pool 1D kernel for k=2, s=2 without output clamp (float16).

Parameters of arm_nn_maxpool1d_k2s2_nhwc_noclip_f16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of channels.
in_wint32_tinInput width. Currently unused by the kernel.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][in_c]`.
out_wint32_tinOutput width. Output position `ow` reads input positions `2*ow..2*ow+1`.
function

Specialized NHWC max-pool 1D kernel for k=2, s=2 with clamp (float16).

Include/arm_nnsupportfunctions_flt.h:1249

void arm_nn_maxpool1d_k2s2_nhwc_f16(
const float16_t *x_nhwc,
int32_t in_c,
int32_t in_w,
float16_t *out,
int32_t out_w,
float16_t act_min,
float16_t act_max
)

Specialized NHWC max-pool 1D kernel for k=2, s=2 with clamp (float16).

Parameters of arm_nn_maxpool1d_k2s2_nhwc_f16
NameTypeDirectionDescription
x_nhwcconst float16_t *inInput row in NHWC layout with shape `[in_w][in_c]`.
in_cint32_tinNumber of channels.
in_wint32_tinInput width. Currently unused by the kernel.
outfloat16_t *outOutput row in NHWC layout with shape `[out_w][in_c]`.
out_wint32_tinOutput width. Output position `ow` reads input positions `2*ow..2*ow+1`.
act_minfloat16_tinLower clamp bound applied to `out`.
act_maxfloat16_tinUpper clamp bound applied to `out`.
function

Matrix multiply with non-transposed lhs and transposed rhs rows (float32).

Include/arm_nnsupportfunctions_flt.h:1273

arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f16(
const float16_t *lhs,
const float16_t *rhs,
const float16_t *bias,
float16_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float16_t activation_min,
float16_t activation_max
)

Matrix multiply with non-transposed lhs and transposed rhs rows (float32).

Parameters of arm_nn_mat_mult_nt_t_f16
NameTypeDirectionDescription
lhsconst float16_t *inLeft-hand matrix stored row-major.
rhsconst float16_t *inRight-hand matrix stored row-major, one row per output channel.
biasconst float16_t *inOptional bias vector.
dstfloat16_t *outOutput matrix.
lhs_rowsint32_tinNumber of rows in `lhs`.
rhs_rowsint32_tinNumber of rows in `rhs`.
rhs_colsint32_tinNumber of columns in `rhs`.
row_address_offsetint32_tinOutput row stride, expressed in elements.
activation_minfloat16_tinLower clamp bound.
activation_maxfloat16_tinUpper clamp bound.
Returns of arm_nn_mat_mult_nt_t_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

armnnmatmultnttf16 with every MVE accumulator lane in float16 (no blockwise fold).

Include/arm_nnsupportfunctions_flt.h:1301

arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f16_acc16(
const float16_t *lhs,
const float16_t *rhs,
const float16_t *bias,
float16_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float16_t activation_min,
float16_t activation_max
)

arm_nn_mat_mult_nt_t_f16 with every MVE accumulator lane in float16 (no blockwise fold).

Same arguments, return codes and scalar leg as arm_nn_mat_mult_nt_t_f16; see its accumulation note.

Parameters of arm_nn_mat_mult_nt_t_f16_acc16
NameTypeDirectionDescription
lhsconst float16_t *inLeft-hand matrix, row-major `[lhs_rows, rhs_cols]`.
rhsconst float16_t *inRight-hand matrix, row-major `[rhs_rows, rhs_cols]` (transposed operand).
biasconst float16_t *inOptional bias vector of `rhs_rows` elements.
dstfloat16_t *outOutput matrix.
lhs_rowsint32_tinNumber of rows in `lhs`.
rhs_rowsint32_tinNumber of rows in `rhs`.
rhs_colsint32_tinShared reduction dimension `K`.
row_address_offsetint32_tinOutput row stride, expressed in elements.
activation_minfloat16_tinLower clamp bound.
activation_maxfloat16_tinUpper clamp bound.
Returns of arm_nn_mat_mult_nt_t_f16_acc16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Matrix multiply with non-transposed lhs and packed non-transposed rhs (float16).

Include/arm_nnsupportfunctions_flt.h:1342

arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f16(
const float16_t *lhs,
const float16_t *rhs_packed,
const float16_t *bias,
float16_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float16_t activation_min,
float16_t activation_max
)

Matrix multiply with non-transposed lhs and packed non-transposed rhs (float16).

Parameters of arm_nn_mat_mult_nt_n_packed_f16
NameTypeDirectionDescription
lhsconst float16_t *inLeft-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`.
rhs_packedconst float16_t *inRight-hand matrix with logical shape `[rhs_cols, rhs_rows]`, packed in column blocks of 8. The final block uses the same packed stride and inactive tail lanes are ignored.
biasconst float16_t *inOptional bias vector.
dstfloat16_t *outOutput matrix.
lhs_rowsint32_tinNumber of rows in `lhs`.
rhs_rowsint32_tinNumber of logical output columns in the unpacked rhs matrix.
rhs_colsint32_tinShared reduction dimension `K`.
row_address_offsetint32_tinOutput row stride, expressed in elements.
activation_minfloat16_tinLower clamp bound.
activation_maxfloat16_tinUpper clamp bound.
Returns of arm_nn_mat_mult_nt_n_packed_f16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

armnnmatmultntnpackedf16 with every MVE accumulator lane in float16 (no blockwise fold).

Include/arm_nnsupportfunctions_flt.h:1370

arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f16_acc16(
const float16_t *lhs,
const float16_t *rhs_packed,
const float16_t *bias,
float16_t *dst,
int32_t lhs_rows,
int32_t rhs_rows,
int32_t rhs_cols,
int32_t row_address_offset,
float16_t activation_min,
float16_t activation_max
)

arm_nn_mat_mult_nt_n_packed_f16 with every MVE accumulator lane in float16 (no blockwise fold).

Same arguments, return codes and scalar leg as arm_nn_mat_mult_nt_n_packed_f16; see its accumulation note.

Parameters of arm_nn_mat_mult_nt_n_packed_f16_acc16
NameTypeDirectionDescription
lhsconst float16_t *inLeft-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`.
rhs_packedconst float16_t *inRight-hand matrix packed in column blocks of 8.
biasconst float16_t *inOptional bias vector.
dstfloat16_t *outOutput matrix.
lhs_rowsint32_tinNumber of rows in `lhs`.
rhs_rowsint32_tinNumber of logical output columns in the unpacked rhs matrix.
rhs_colsint32_tinShared reduction dimension `K`.
row_address_offsetint32_tinOutput row stride, expressed in elements.
activation_minfloat16_tinLower clamp bound.
activation_maxfloat16_tinUpper clamp bound.
Returns of arm_nn_mat_mult_nt_n_packed_f16_acc16
Description
`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments.
function

Update LSTM function for an iteration step using float16 input, output and state.

Include/arm_nnsupportfunctions_flt.h:1394

arm_cmsis_nn_status arm_nn_lstm_step_f16(
const float16_t *data_in,
const float16_t *hidden_in,
float16_t *hidden_out,
const cmsis_nn_lstm_params_f16 *params,
cmsis_nn_lstm_context_f16 *buffers,
const int32_t batch_offset
)

Update LSTM function for an iteration step using float16 input, output and state.

Parameters of arm_nn_lstm_step_f16
NameTypeDirectionDescription
data_inconst float16_t *inData input pointer.
hidden_inconst float16_t *inHidden state / recurrent input pointer. May be NULL for the first step.
hidden_outfloat16_t *outHidden state / recurrent output pointer.
paramsconst cmsis_nn_lstm_params_f16 *inStruct containing all information about the LSTM operator.
bufferscmsis_nn_lstm_context_f16 *in, outStruct containing pointers to mutable cell-state storage.
batch_offsetconst int32_tinNumber of timesteps between consecutive batches.
Returns of arm_nn_lstm_step_f16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params/buffers or buffers->cell_state, batch_offset <= 0).
function

Update GRU function for a single iteration step using float16 data.

Include/arm_nnsupportfunctions_flt.h:1414

arm_cmsis_nn_status arm_nn_gru_step_f16(
const float16_t *data_in,
const float16_t *hidden_in,
float16_t *hidden_out,
const cmsis_nn_gru_params_f16 *params,
cmsis_nn_gru_context_f16 *buffers,
const int32_t batch_offset
)

Update GRU function for a single iteration step using float16 data.

Parameters of arm_nn_gru_step_f16
NameTypeDirectionDescription
data_inconst float16_t *inData input pointer for this time step.
hidden_inconst float16_t *inRecurrent input pointer. NULL for the first step (h_prev = 0).
hidden_outfloat16_t *outHidden-state output pointer for this time step.
paramsconst cmsis_nn_gru_params_f16 *inStruct describing the GRU operator.
bufferscmsis_nn_gru_context_f16 *in, outScratch buffers. temp1 (>= hidden_size) is required when reset_after == 0.
batch_offsetconst int32_tinNumber of timesteps between consecutive batches.
Returns of arm_nn_gru_step_f16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params, batch_offset <= 0, or missing temp1 when reset_after == 0).
function

Pack a single convolution patch into one row of a contiguous float32 patch matrix.

Include/arm_nnsupportfunctions_flt.h:1424

void arm_nn_pack_conv_patch_f16(
const float16_t *input,
int32_t in_h,
int32_t in_w,
int32_t in_c,
int32_t kernel_h,
int32_t kernel_w,
int32_t stride_h,
int32_t stride_w,
int32_t pad_h,
int32_t pad_w,
int32_t dilation_h,
int32_t dilation_w,
int32_t out_y,
int32_t out_x,
float16_t pad_value,
float16_t *patch_row
)

Pack a single convolution patch into one row of a contiguous float32 patch matrix.

Developers familiar with im2row/im2col terminology can think of this as packing one output patch into one row.

Parameters of arm_nn_pack_conv_patch_f16
NameTypeDirectionDescription
inputconst float16_t *inInput tensor for one batch in NHWC layout with shape `[in_h][in_w][in_c]`.
in_hint32_tinInput height.
in_wint32_tinInput width.
in_cint32_tinNumber of input channels.
kernel_hint32_tinKernel height.
kernel_wint32_tinKernel width.
stride_hint32_tinVertical stride.
stride_wint32_tinHorizontal stride.
pad_hint32_tinTop padding.
pad_wint32_tinLeft padding.
dilation_hint32_tinVertical dilation.
dilation_wint32_tinHorizontal dilation.
out_yint32_tinOutput row index of the patch to pack.
out_xint32_tinOutput column index of the patch to pack.
pad_valuefloat16_tinValue written for taps that fall outside the input.
patch_rowfloat16_t *outDestination row of `kernel_h * kernel_w * in_c` elements, ordered `[kernel_h][kernel_w][in_c]`.
function

Specialized softmax helper for a single float16 row of length 2.

Include/arm_nnsupportfunctions_flt.h:1447

void arm_nn_softmax_1x2_f16(const float16_t *in, float16_t *out)

Specialized softmax helper for a single float16 row of length 2.

Parameters of arm_nn_softmax_1x2_f16
NameTypeDirectionDescription
inconst float16_t *inPointer to two contiguous float16 input values.
outfloat16_t *outPointer to two contiguous float16 output values.
function

Update LSTM function for an iteration step using float32 input, output and state.

Include/arm_nnsupportfunctions_flt.h:1466

arm_cmsis_nn_status arm_nn_lstm_step_f32(
const float32_t *data_in,
const float32_t *hidden_in,
float32_t *hidden_out,
const cmsis_nn_lstm_params_f32 *params,
cmsis_nn_lstm_context_f32 *buffers,
const int32_t batch_offset
)

Update LSTM function for an iteration step using float32 input, output and state.

Parameters of arm_nn_lstm_step_f32
NameTypeDirectionDescription
data_inconst float32_t *inData input pointer.
hidden_inconst float32_t *inHidden state / recurrent input pointer. May be NULL for the first step.
hidden_outfloat32_t *outHidden state / recurrent output pointer.
paramsconst cmsis_nn_lstm_params_f32 *inStruct containing all information about the LSTM operator.
bufferscmsis_nn_lstm_context_f32 *in, outStruct containing pointers to mutable cell-state storage.
batch_offsetconst int32_tinNumber of timesteps between consecutive batches.
Returns of arm_nn_lstm_step_f32
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params/buffers or buffers->cell_state, batch_offset <= 0).
function

Update GRU function for a single iteration step using float32 data.

Include/arm_nnsupportfunctions_flt.h:1486

arm_cmsis_nn_status arm_nn_gru_step_f32(
const float32_t *data_in,
const float32_t *hidden_in,
float32_t *hidden_out,
const cmsis_nn_gru_params_f32 *params,
cmsis_nn_gru_context_f32 *buffers,
const int32_t batch_offset
)

Update GRU function for a single iteration step using float32 data.

Parameters of arm_nn_gru_step_f32
NameTypeDirectionDescription
data_inconst float32_t *inData input pointer for this time step.
hidden_inconst float32_t *inRecurrent input pointer. NULL for the first step (h_prev = 0).
hidden_outfloat32_t *outHidden-state output pointer for this time step.
paramsconst cmsis_nn_gru_params_f32 *inStruct describing the GRU operator.
bufferscmsis_nn_gru_context_f32 *in, outScratch buffers. temp1 (>= hidden_size) is required when reset_after == 0.
batch_offsetconst int32_tinNumber of timesteps between consecutive batches.
Returns of arm_nn_gru_step_f32
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params, batch_offset <= 0, or missing temp1 when reset_after == 0).