Skip to content
heliaCORE
API reference
HELIA HUB

arm_nnsupportfunctions.h

Machine-readable model

macro

Include/arm_nnsupportfunctions.h:45

#define USE_FAST_DW_CONV_S16_FUNCTION(dw_conv_params, filter_dims, input_dims, output_dims) (dw_conv_params->ch_mult == 1 && \ arm_nn_dw_conv_opt_dilation_supported(dw_conv_params, input_dims, filter_dims, output_dims) && \ filter_dims->w * filter_dims->h < 512)
Parameters of USE_FAST_DW_CONV_S16_FUNCTION
NameDescription
dw_conv_params
filter_dims
input_dims
output_dims
function

Minimum of two scalar f16 values.

Include/arm_nnsupportfunctions.h:98

static _Float16 arm_nn_min_f16h(_Float16 a, _Float16 b)

Minimum of two scalar f16 values.

With ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 minNum via VMINNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MIN, an ordered compare, so its NaN handling depends on operand order: a NaN b is returned, a NaN a is not. Do not rely on NaN suppression on non-MVE builds.

Parameters of arm_nn_min_f16h
NameTypeDirectionDescription
a_Float16inFirst operand
b_Float16inSecond operand
Returns of arm_nn_min_f16h
Description
The smaller of `a` and `b`
function

Maximum of two scalar f16 values.

Include/arm_nnsupportfunctions.h:121

static _Float16 arm_nn_max_f16h(_Float16 a, _Float16 b)

Maximum of two scalar f16 values.

With ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 maxNum via VMAXNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MAX, an ordered compare, so its NaN handling depends on operand order: a NaN b is returned, a NaN a is not. Do not rely on NaN suppression on non-MVE builds.

Parameters of arm_nn_max_f16h
NameTypeDirectionDescription
a_Float16inFirst operand
b_Float16inSecond operand
Returns of arm_nn_max_f16h
Description
The larger of `a` and `b`
function

Returns x when x is NaN, otherwise y.

Include/arm_nnsupportfunctions.h:167

static _Float16 arm_nn_propagate_nan_f16h(_Float16 x, _Float16 y)

Returns x when x is NaN, otherwise y.

Both the NaN test and the select are performed on the bit patterns: the test is (bits & 0x7FFF) > 0x7C00 (all-ones exponent, non-zero mantissa), which is integer arithmetic that -ffinite-math-only (implied by the shipped -Ofast) has no license to fold, unlike the former floating-point self-compare x != x (#333 / #334); the bit-pattern select neither expands to an HFmode conditional move (PR target/118460) nor quiets/retags the NaN payload. This helper backs the f16 elementwise clamp and, via arm_nn_clamp_scalar_f16 / arm_nn_clamp_propagate_nan_f16h, the other f16 scalar clamp users arm_svdf_f16, arm_max_pool_f16 / arm_avg_pool_f16, the packed f16 matmul (arm_nn_mat_mult_nt_n_packed_f16), the scalar f16 RELU/RELU6/LEAKY_RELU activation legs, and arm_nn_vector_clamp_f16’s scalar leg (conv/depthwise/transpose-conv f16, the 3x3 depthwise, and arm_nn_maxpool1d_f16) so wherever that scalar clamp runs, a NaN passes through it at every optimization level on the gated toolchains. That is a guarantee about the clamp, not the whole kernel: which builds run the scalar clamp, and whether a NaN survives the rest of the kernel to reach it, is per kernel several of these callers clamp with vmaxnmq/vminnmq on MVE builds (a NaN resolves to a bound there), and arm_max_pool_f16’s max reduction drops a NaN before the clamp. The kernels with a NaN

Parameters of arm_nn_propagate_nan_f16h
NameTypeDirectionDescription
x_Float16inValue whose NaN-ness selects the result. Returned unchanged when it is NaN.
y_Float16inValue returned when `x` is not NaN
Returns of arm_nn_propagate_nan_f16h
Description
`x` if `x` is NaN, otherwise `y`
function

Drop-in equivalent of ARMNNCLAMP(x, h, l) for scalar Float16 operands.

Include/arm_nnsupportfunctions.h:193

static _Float16 arm_nn_clamp_f16h(_Float16 x, _Float16 h, _Float16 l)

Drop-in equivalent of ARM_NN_CLAMP(x, h, l) for scalar _Float16 operands.

Includes the macro’s NaN behaviour: ARM_NN_MIN(NaN, h) is h, so a NaN input resolves to the high bound, exactly as the macro does. Use arm_nn_clamp_propagate_nan_f16h() where TFLite NaN propagation is required.

Parameters of arm_nn_clamp_f16h
NameTypeDirectionDescription
x_Float16inValue to clamp
h_Float16inUpper bound
l_Float16inLower bound
Returns of arm_nn_clamp_f16h
Description
`x` clamped to [`l`, `h`]
function

Scalar f16 clamp with TFLite NaN semantics: NaN passes through unchanged.

Include/arm_nnsupportfunctions.h:212

static _Float16 arm_nn_clamp_propagate_nan_f16h(_Float16 x, _Float16 l, _Float16 h)

Scalar f16 clamp with TFLite NaN semantics: NaN passes through unchanged.

Mirrors the MVE idiom in arm_nn_clamp_propagate_nan_mve_f16() (lower bound first, then upper bound, then restore NaN lanes). The NaN restore in arm_nn_propagate_nan_f16h() tests the integer bit pattern, so it holds at every optimization level including the shipped -Ofast; see #333 / #334. Bounds are assumed ordered (l <= h); inverted bounds are unspecified.

Parameters of arm_nn_clamp_propagate_nan_f16h
NameTypeDirectionDescription
x_Float16inValue to clamp
l_Float16inLower bound
h_Float16inUpper bound
Returns of arm_nn_clamp_propagate_nan_f16h
Description
`x` clamped to [`l`, `h`], or `x` itself when it is NaN
function

Absolute value of a scalar f16 value.

Include/arm_nnsupportfunctions.h:224

static _Float16 arm_nn_abs_f16h(_Float16 x)

Absolute value of a scalar f16 value.

Parameters of arm_nn_abs_f16h
NameTypeDirectionDescription
x_Float16inInput value
Returns of arm_nn_abs_f16h
Description
|`x`|
function

Fold one dimension into a running buffer-size product, reporting overflow as -1.

Include/arm_nnsupportfunctions.h:307

static int64_t arm_nn_size_mul(const int64_t acc, const int64_t factor)

Fold one dimension into a running buffer-size product, reporting overflow as -1.

Buffer-size queries return an int32_t byte count, so the product of the dimensions they multiply has to be rejected as soon as it cannot fit. Folding one factor at a time keeps the accumulator bounded: an accumulator already known to be <= INT32_MAX times a factor <= INT32_MAX cannot exceed about 2^62, so the int64_t accumulator itself never wraps. Chaining raw (int64_t) casts across three or more int32_t dims does not have that property - 65536 * 65536 * 65536 * 65536 is exactly 2^64 and folds back to 0, which would sail through a trailing “> INT32_MAX” test.

Parameters of arm_nn_size_mul
NameTypeDirectionDescription
accconst int64_tinRunning product, or -1 if an earlier fold already overflowed.
factorconst int64_tinNext factor to fold in.
Returns of arm_nn_size_mul
Description
acc * factor, or -1 if acc is already -1, factor is negative or out of int32_t range, or the product exceeds INT32_MAX.
function

Add to a running buffer-size product, reporting overflow as -1.

Include/arm_nnsupportfunctions.h:332

static int64_t arm_nn_size_add(const int64_t acc, const int64_t addend)

Add to a running buffer-size product, reporting overflow as -1.

Companion to arm_nn_size_mul() for the sizers that append a fixed slack term.

Parameters of arm_nn_size_add
NameTypeDirectionDescription
accconst int64_tinRunning product, or -1 if an earlier step already overflowed.
addendconst int64_tinValue to add. Must be non-negative.
Returns of arm_nn_size_add
Description
acc + addend, or -1 if acc is already -1 or the sum exceeds INT32_MAX.
macro

definition to pack four 8 bit values.

Include/arm_nnsupportfunctions.h:358

#define PACK_S8x4_32x1(v0, v1, v2, v3) ((int32_t)((((uint32_t)(v0)) & 0xFFu) | ((((uint32_t)(v1)) & 0xFFu) << 8) | ((((uint32_t)(v2)) & 0xFFu) << 16) | \ ((((uint32_t)(v3)) & 0xFFu) << 24)))

definition to pack four 8 bit values.

Byte lanes are masked and shifted in uint32_t so a negative value never feeds a signed left shift (UB); masking before the shift keeps the same bits the old shift-then-mask form kept. Bit-identical for every input. Deliberate divergence from upstream ARM-software/CMSIS-NN, which still carries the signed-shift form do not paste the upstream text back on a sync (issue #357).

Parameters of PACK_S8x4_32x1
NameDescription
v0
v1
v2
v3
macro

definition to pack two 16 bit values.

Include/arm_nnsupportfunctions.h:369

#define PACK_Q15x2_32x1(v0, v1) ((int32_t)((((uint32_t)(v0)) & 0xFFFFu) | (((uint32_t)(v1)) << 16)))

definition to pack two 16 bit values.

Same treatment: the high half is shifted in uint32_t, not int32_t, so a negative v1 is defined; the low half keeps its mask. Bit-identical for every input. Same deliberate upstream divergence as PACK_S8x4_32x1 above.

Parameters of PACK_Q15x2_32x1
NameDescription
v0
v1
function

Map an output index to the nearest input index for resize.

Include/arm_nnsupportfunctions.h:391

static int32_t GetNearestNeighbor(
const int input_value,
const int32_t input_size,
const float scale,
const float offset,
const bool align_corners,
const bool half_pixel_centers
)

Map an output index to the nearest input index for resize.

This helper follows the TensorFlow Lite nearest-neighbor resize mapping rules.

Parameters of GetNearestNeighbor
NameTypeDirectionDescription
input_valueconst intinOutput index (x or y).
input_sizeconst int32_tinInput size along the same axis.
scaleconst floatinPrecomputed scaling factor for the axis.
offsetconst floatinPrecomputed offset for the axis.
align_cornersconst boolinIf true, use align-corners scaling.
half_pixel_centersconst boolinIf true, use half-pixel center offset.
Returns of GetNearestNeighbor
Description
Nearest input index for the given output index.
function

Check if convolution parameters correspond to a 1x1 convolution.

Include/arm_nnsupportfunctions.h:416

static bool arm_nn_is_convolve_1x1(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims
)

Check if convolution parameters correspond to a 1x1 convolution.

Parameters of arm_nn_is_convolve_1x1
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters
input_dimsconst cmsis_nn_dims *inInput dimensions
filter_dimsconst cmsis_nn_dims *inFilter dimensions
Returns of arm_nn_is_convolve_1x1
Description
true if parameters describe a 1x1 convolution, false otherwise.
function

Check if a 1x1 convolution qualifies for the fast (unit stride) path.

Include/arm_nnsupportfunctions.h:432

static bool arm_nn_is_convolve_1x1_fast(const cmsis_nn_conv_params *conv_params)

Check if a 1x1 convolution qualifies for the fast (unit stride) path.

Parameters of arm_nn_is_convolve_1x1_fast
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters
Returns of arm_nn_is_convolve_1x1_fast
Description
true if stride is 1x1, false otherwise.
function

Check if convolution parameters correspond to a 1xN convolution.

Include/arm_nnsupportfunctions.h:444

static bool arm_nn_is_convolve_1_x_n(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims
)

Check if convolution parameters correspond to a 1xN convolution.

Parameters of arm_nn_is_convolve_1_x_n
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters
input_dimsconst cmsis_nn_dims *inInput dimensions
filter_dimsconst cmsis_nn_dims *inFilter dimensions
Returns of arm_nn_is_convolve_1_x_n
Description
true if parameters describe a 1xN convolution, false otherwise.
function

Check that armconvolve1xns4() handles the horizontal padding of a 1xN convolution.

Include/arm_nnsupportfunctions.h:473

static bool arm_nn_convolve_1_x_n_padding_supported(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Check that arm_convolve_1_x_n_s4() handles the horizontal padding of a 1xN convolution.

The kernel places pad.w columns on the left and pad.w + (total_pad % 2) on the right, where total_pad = (output W - 1) * stride.w + filter W - input W, and needs the output columns that read padding to fit in output W. Its padded-column code also assumes that each such column reads at least one input column and that the filter is no wider than the input; otherwise it forms input and filter addresses outside the tensors. On MVE builds another pad placement, or too many padded columns, returns ARM_CMSIS_NN_FAILURE. A VALID layer whose stride leaves trailing input unused (negative total_pad) is therefore rejected, and the wrapper routes it to another convolution. The kernel also computes a single output row, so vertical padding or an output height other than 1 is rejected. A non-positive stride.w is left to the kernel’s argument checks.

Parameters of arm_nn_convolve_1_x_n_padding_supported
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters
input_dimsconst cmsis_nn_dims *inInput dimensions
filter_dimsconst cmsis_nn_dims *inFilter dimensions
output_dimsconst cmsis_nn_dims *inOutput dimensions
Returns of arm_nn_convolve_1_x_n_padding_supported
Description
true when `arm_convolve_1_x_n_s4()` handles the padding, false otherwise.
function

Check that armconvolve1xns8() accepts the padding and output shape of a 1xN convolution.

Include/arm_nnsupportfunctions.h:517

static bool arm_nn_convolve_1_x_n_s8_padding_supported(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Check that arm_convolve_1_x_n_s8() accepts the padding and output shape of a 1xN convolution.

The kernel computes a single output row for any pad.w >= 0 and any output width, including an odd total padding, a filter wider than the input and a VALID layer whose stride leaves trailing input unused. It rejects vertical padding, an output height other than 1, a negative pad.w and an empty filter; the wrapper routes those layers to another convolution. A non-positive stride.w is left to the kernel’s argument checks.

Parameters of arm_nn_convolve_1_x_n_s8_padding_supported
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters
filter_dimsconst cmsis_nn_dims *inFilter dimensions
output_dimsconst cmsis_nn_dims *inOutput dimensions
Returns of arm_nn_convolve_1_x_n_s8_padding_supported
Description
true when `arm_convolve_1_x_n_s8()` computes the layer, false otherwise.
function

Count the output columns of a 1xN convolution whose window reads padding.

Include/arm_nnsupportfunctions.h:539

static void arm_nn_convolve_1_x_n_padded_columns(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int64_t *left_num,
int64_t *right_num
)

Count the output columns of a 1xN convolution whose window reads padding.

Output column j reads input columns j * stride.w - pad.w to j * stride.w - pad.w + filter W - 1. The leading columns whose window starts before the input are left-padded; of the others, the trailing columns whose window ends past the input are right-padded. A window can do both only when it is left-padded.

Parameters of arm_nn_convolve_1_x_n_padded_columns
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters. stride.w >= 1 and pad.w >= 0.
input_dimsconst cmsis_nn_dims *inInput dimensions. w >= 0.
filter_dimsconst cmsis_nn_dims *inFilter dimensions. w >= 1.
output_dimsconst cmsis_nn_dims *inOutput dimensions. w >= 0.
left_numint64_t *outNumber of left-padded output columns, output W at most.
right_numint64_t *outNumber of right-padded output columns, output W - left_num at most.
function

Check if the dilation, stride and padding of a depthwise layer allow the armdepthwiseconvs8opt() or armdepthwiseconvfasts16() route.

Include/arm_nnsupportfunctions.h:572

static bool arm_nn_dw_conv_opt_dilation_supported(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Check if the dilation, stride and padding of a depthwise layer allow the arm_depthwise_conv_s8_opt() or arm_depthwise_conv_fast_s16() route.

Parameters of arm_nn_dw_conv_opt_dilation_supported
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters
input_dimsconst cmsis_nn_dims *inInput dimensions
filter_dimsconst cmsis_nn_dims *inFilter dimensions
output_dimsconst cmsis_nn_dims *inOutput dimensions
Returns of arm_nn_dw_conv_opt_dilation_supported
Description
true for an undilated layer (dilation 1 in both dimensions), or for a 1D layer dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, no vertical padding, dilation.h == 1 and dilation.w >= 1. false otherwise.
function

Converts the elements from a s8 vector to a s16 vector with an added offset.

Include/arm_nnsupportfunctions.h:649

void arm_q7_to_q15_with_offset(const int8_t *src, int16_t *dst, int32_t block_size, int16_t offset)

Converts the elements from a s8 vector to a s16 vector with an added offset.

Output elements are ordered. The equation used for the conversion process is:

dst[n] = (int16_t) src[n] + offset; 0 <= n < block_size.

Parameters of arm_q7_to_q15_with_offset
NameTypeDirectionDescription
srcconst int8_t *inpointer to the s8 input vector
dstint16_t *outpointer to the s16 output vector
block_sizeint32_tinlength of the input vector
offsetint16_tins16 offset to be added to each input vector element.
function

Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel.

Include/arm_nnsupportfunctions.h:695

int32_t arm_depthwise_conv_s8_opt_get_buffer_size_mve(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with MVE extension.

The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.

Parameters of arm_depthwise_conv_s8_opt_get_buffer_size_mve
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_s8_opt_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel.

Include/arm_nnsupportfunctions.h:710

int32_t arm_depthwise_conv_s8_opt_get_buffer_size_dsp(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with DSP extension.

The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.

Parameters of arm_depthwise_conv_s8_opt_get_buffer_size_dsp
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_s8_opt_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Depthwise conv on an im2col buffer where the input channel equals output channel.

Include/arm_nnsupportfunctions.h:732

int8_t * arm_nn_depthwise_conv_s8_core(
const int8_t *row,
const int16_t *col,
const uint16_t num_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t kernel_size,
const int32_t *const output_bias,
int8_t *out
)

Depthwise conv on an im2col buffer where the input channel equals output channel.

Supported framework: TensorFlow Lite micro.

Parameters of arm_nn_depthwise_conv_s8_core
NameTypeDirectionDescription
rowconst int8_t *inpointer to row
colconst int16_t *inpointer to im2col buffer, always consists of 2 columns.
num_chconst uint16_tinnumber of channels
out_shiftconst int32_t *inpointer to per output channel requantization shift parameter.
out_multconst int32_t *inpointer to per output channel requantization multiplier parameter.
out_offsetconst int32_tinoutput tensor offset.
activation_minconst int32_tinminimum value to clamp the output to. Range : int8
activation_maxconst int32_tinmaximum value to clamp the output to. Range : int8
kernel_sizeconst uint16_tinnumber of elements in one column.
output_biasconst int32_t *constinper output channel bias. Range : int32
outint8_t *outpointer to output
Returns of arm_nn_depthwise_conv_s8_core
Description
The function returns one of the two 1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available.
function

General Matrix-multiplication function with per-channel requantization.

Include/arm_nnsupportfunctions.h:766

int8_t * arm_nn_mat_mult_s8(
const int8_t *input_row,
const int8_t *input_col,
const uint16_t output_ch,
const uint16_t col_batches,
const int32_t *output_shift,
const int32_t *output_mult,
const int32_t out_offset,
const int32_t col_offset,
const int32_t row_offset,
const int16_t out_activation_min,
const int16_t out_activation_max,
const uint16_t row_len,
const int32_t *const bias,
int8_t *out
)

General Matrix-multiplication function with per-channel requantization.

Supported framework: TensorFlow Lite

Parameters of arm_nn_mat_mult_s8
NameTypeDirectionDescription
input_rowconst int8_t *inpointer to row operand
input_colconst int8_t *inpointer to col operand
output_chconst uint16_tinnumber of rows of input_row
col_batchesconst uint16_tinnumber of column batches. Range: 1 to 4
output_shiftconst int32_t *inpointer to per output channel requantization shift parameter.
output_multconst int32_t *inpointer to per output channel requantization multiplier parameter.
out_offsetconst int32_tinoutput tensor offset.
col_offsetconst int32_tininput tensor(col) offset.
row_offsetconst int32_tinkernel offset(row). Not used.
out_activation_minconst int16_tinminimum value to clamp the output to. Range : int8
out_activation_maxconst int16_tinmaximum value to clamp the output to. Range : int8
row_lenconst uint16_tinnumber of elements in each row
biasconst int32_t *constinper output channel bias. Range : int32
outint8_t *in, outpointer to output
Returns of arm_nn_mat_mult_s8
Description
The function returns one of the two 1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available.
function

Matrix-multiplication function for convolution with per-channel requantization for 16 bits convolution.

Include/arm_nnsupportfunctions.h:804

int16_t * arm_nn_mat_mult_kernel_s16(
const int8_t *input_a,
const int16_t *input_b,
const int32_t output_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t activation_min,
const int32_t activation_max,
const int32_t num_col_a,
const cmsis_nn_bias_data *const bias_data,
int16_t *out_0,
const int32_t row_address_offset
)

Matrix-multiplication function for convolution with per-channel requantization for 16 bits convolution.

This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.

Parameters of arm_nn_mat_mult_kernel_s16
NameTypeDirectionDescription
input_aconst int8_t *inpointer to operand A
input_bconst int16_t *inpointer to operand B, always consists of 2 vectors.
output_chconst int32_tinnumber of rows of A
out_shiftconst int32_t *inpointer to per output channel requantization shift parameter.
out_multconst int32_t *inpointer to per output channel requantization multiplier parameter.
activation_minconst int32_tinminimum value to clamp the output to. Range : int16
activation_maxconst int32_tinmaximum value to clamp the output to. Range : int16
num_col_aconst int32_tinnumber of columns of A
bias_dataconst cmsis_nn_bias_data *constinpointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct.
out_0int16_t *in, outpointer to output
row_address_offsetconst int32_tinAddress offset between rows in output.
Returns of arm_nn_mat_mult_kernel_s16
Description
The function returns one of the two 1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available.
function

General Vector by Matrix multiplication with requantization and storage of result.

Include/arm_nnsupportfunctions.h:843

arm_cmsis_nn_status arm_nn_mat_mul_core_1x_s8(
int32_t row_elements,
const int32_t skipped_row_elements,
const int8_t *row_base_ref,
const int8_t *col_base_ref,
const int32_t out_ch,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const int32_t *bias,
int8_t *output
)

General Vector by Matrix multiplication with requantization and storage of result.

Pseudo-code *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in ‘output’

Parameters of arm_nn_mat_mul_core_1x_s8
NameTypeDirectionDescription
row_elementsint32_tinnumber of row elements
skipped_row_elementsconst int32_tinnumber of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch
row_base_refconst int8_t *inpointer to row operand
col_base_refconst int8_t *inpointer to col operand
out_chconst int32_tinNumber of output channels
conv_paramsconst cmsis_nn_conv_params *inPointer to convolution parameters like offsets and activation values
quant_paramsconst cmsis_nn_per_channel_quant_params *inPointer to per-channel quantization parameters
biasconst int32_t *inPointer to optional per-channel bias
outputint8_t *outPointer to output where int8 results are stored.
Returns of arm_nn_mat_mul_core_1x_s8
Description
The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory.
function

General Vector by Matrix multiplication with requantization, storage of result and int4 weights packed into an int8 buffer.

Include/arm_nnsupportfunctions.h:881

arm_cmsis_nn_status arm_nn_mat_mul_core_1x_s4(
int32_t row_elements,
const int32_t skipped_row_elements,
const int8_t *row_base_ref,
const int8_t *col_base_ref,
const int32_t out_ch,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const int32_t *bias,
int8_t *output
)

General Vector by Matrix multiplication with requantization, storage of result and int4 weights packed into an int8 buffer.

Pseudo-code as int8 example. Int4 filter data will be unpacked. *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in ‘output’

Parameters of arm_nn_mat_mul_core_1x_s4
NameTypeDirectionDescription
row_elementsint32_tinnumber of row elements
skipped_row_elementsconst int32_tinnumber of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch
row_base_refconst int8_t *inpointer to row operand
col_base_refconst int8_t *inpointer to col operand as packed int4
out_chconst int32_tinNumber of output channels
conv_paramsconst cmsis_nn_conv_params *inPointer to convolution parameters like offsets and activation values
quant_paramsconst cmsis_nn_per_channel_quant_params *inPointer to per-channel quantization parameters
biasconst int32_t *inPointer to optional per-channel bias
outputint8_t *outPointer to output where int8 results are stored.
Returns of arm_nn_mat_mul_core_1x_s4
Description
The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory.
function

Matrix-multiplication with requantization & activation function for four rows and one column.

Include/arm_nnsupportfunctions.h:908

int8_t * arm_nn_mat_mul_core_4x_s8(
const int32_t row_elements,
const int32_t offset,
const int8_t *row_base,
const int8_t *col_base,
const int32_t out_ch,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const int32_t *bias,
int8_t *output
)

Matrix-multiplication with requantization & activation function for four rows and one column.

Compliant to TFLM int8 specification. MVE implementation only

Parameters of arm_nn_mat_mul_core_4x_s8
NameTypeDirectionDescription
row_elementsconst int32_tinnumber of row elements
offsetconst int32_tinoffset between rows. Can be the same as row_elements. For e.g, in a 1x1 conv scenario with stride as 1.
row_baseconst int8_t *inpointer to row operand
col_baseconst int8_t *inpointer to col operand
out_chconst int32_tinNumber of output channels
conv_paramsconst cmsis_nn_conv_params *inPointer to convolution parameters like offsets and activation values
quant_paramsconst cmsis_nn_per_channel_quant_params *inPointer to per-channel quantization parameters
biasconst int32_t *inPointer to per-channel bias
outputint8_t *outPointer to output where int8 results are stored.
Returns of arm_nn_mat_mul_core_4x_s8
Description
The function returns the updated output pointer or NULL if implementation is not available.
function

General Matrix-multiplication function with per-channel requantization.

Include/arm_nnsupportfunctions.h:950

arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s4(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *bias,
int8_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t lhs_cols_offset
)

General Matrix-multiplication function with per-channel requantization. This function assumes:

  • LHS input matrix NOT transposed (nt)
  • RHS input matrix transposed (t)
  • RHS is int8 packed with 2x int4
  • LHS is int8
Parameters of arm_nn_mat_mult_nt_t_s4
NameTypeDirectionDescription
lhsconst int8_t *inPointer to the LHS input matrix
rhsconst int8_t *inPointer to the RHS input matrix
biasconst int32_t *inPointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)
dstint8_t *outPointer to the output matrix with "m" rows and "n" columns
dst_multipliersconst int32_t *inPointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shiftsconst int32_t *inPointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rowsconst int32_tinNumber of LHS input rows
rhs_rowsconst int32_tinNumber of RHS input rows
rhs_colsconst int32_tinNumber of LHS/RHS input columns
lhs_offsetconst int32_tinOffset to be applied to the LHS input value
dst_offsetconst int32_tinOffset to be applied the output result
activation_minconst int32_tinMinimum value to clamp down the output. Range : int8
activation_maxconst int32_tinMaximum value to clamp up the output. Range : int8
lhs_cols_offsetconst int32_tinColumn offset between subsequent lhs_rows
Returns of arm_nn_mat_mult_nt_t_s4
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

General Matrix-multiplication function with per-channel requantization.

Include/arm_nnsupportfunctions.h:999

arm_cmsis_nn_status arm_nn_mat_mult_nt_interleaved_t_even_s4(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *bias,
int8_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t lhs_cols_offset
)

General Matrix-multiplication function with per-channel requantization. This function assumes:

  • LHS input matrix NOT transposed (nt)
  • RHS input matrix transposed (t)
  • RHS is int8 packed with 2x int4
  • LHS is int8
  • LHS/RHS input columns must be even numbered
  • LHS must be interleaved. Compare to arm_nn_mat_mult_nt_t_s4 where LHS is not interleaved.
Parameters of arm_nn_mat_mult_nt_interleaved_t_even_s4
NameTypeDirectionDescription
lhsconst int8_t *inPointer to the LHS input matrix
rhsconst int8_t *inPointer to the RHS input matrix
biasconst int32_t *inPointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)
dstint8_t *outPointer to the output matrix with "m" rows and "n" columns
dst_multipliersconst int32_t *inPointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shiftsconst int32_t *inPointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rowsconst int32_tinNumber of LHS input rows
rhs_rowsconst int32_tinNumber of RHS input rows
rhs_colsconst int32_tinNumber of LHS/RHS input columns. Note this must be even.
lhs_offsetconst int32_tinOffset to be applied to the LHS input value
dst_offsetconst int32_tinOffset to be applied the output result
activation_minconst int32_tinMinimum value to clamp down the output. Range : int8
activation_maxconst int32_tinMaximum value to clamp up the output. Range : int8
lhs_cols_offsetconst int32_tinColumn offset between subsequent lhs_rows
Returns of arm_nn_mat_mult_nt_interleaved_t_even_s4
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

General Matrix-multiplication function with per-channel requantization.

Include/arm_nnsupportfunctions.h:1046

arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s8(
const int32_t *weight_sum_buf,
const int8_t *lhs,
const int8_t *rhs,
const int32_t *bias,
int8_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t row_address_offset,
const int32_t lhs_cols_offset
)

General Matrix-multiplication function with per-channel requantization. This function assumes:

  • LHS input matrix NOT transposed (nt)
  • RHS input matrix transposed (t)
Parameters of arm_nn_mat_mult_nt_t_s8
NameTypeDirectionDescription
weight_sum_bufconst int32_t *inPointer to the weight sum multiplied by lhs_offset and summed bias buffer
lhsconst int8_t *inPointer to the LHS input matrix
rhsconst int8_t *inPointer to the RHS input matrix
biasconst int32_t *inPointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)
dstint8_t *outPointer to the output matrix with "m" rows and "n" columns
dst_multipliersconst int32_t *inPointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shiftsconst int32_t *inPointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rowsconst int32_tinNumber of LHS input rows
rhs_rowsconst int32_tinNumber of RHS input rows
rhs_colsconst int32_tinNumber of LHS/RHS input columns
lhs_offsetconst int32_tinOffset to be applied to the LHS input value
dst_offsetconst int32_tinOffset to be applied the output result
activation_minconst int32_tinMinimum value to clamp down the output. Range : int8
activation_maxconst int32_tinMaximum value to clamp up the output. Range : int8
row_address_offsetconst int32_tinAddress offset between rows in output. NOTE: Only used for MVEI extension.
lhs_cols_offsetconst int32_tinColumn offset between subsequent lhs_rows
Returns of arm_nn_mat_mult_nt_t_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

General Matrix-multiplication function with per-channel requantization.

Include/arm_nnsupportfunctions.h:1097

arm_cmsis_nn_status arm_nn_mat_mult_nt_t_1x1_out_s8(
const int32_t *weight_sum_buf,
const int8_t *lhs,
const int8_t *rhs,
const int32_t *bias,
int8_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t row_address_offset,
const int32_t lhs_cols_offset
)

General Matrix-multiplication function with per-channel requantization. Output is calculated with multiple channels in parallel, rather than multiple output indices in a single channel This function assumes:

  • LHS input matrix NOT transposed (nt)
  • RHS input matrix transposed (t)
Parameters of arm_nn_mat_mult_nt_t_1x1_out_s8
NameTypeDirectionDescription
weight_sum_bufconst int32_t *inPointer to the weight sum multiplied by lhs_offset and summed bias buffer
lhsconst int8_t *inPointer to the LHS input matrix
rhsconst int8_t *inPointer to the RHS input matrix
biasconst int32_t *inPointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)
dstint8_t *outPointer to the output matrix with "m" rows and "n" columns
dst_multipliersconst int32_t *inPointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shiftsconst int32_t *inPointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rowsconst int32_tinNumber of LHS input rows
rhs_rowsconst int32_tinNumber of RHS input rows
rhs_colsconst int32_tinNumber of LHS/RHS input columns
lhs_offsetconst int32_tinOffset to be applied to the LHS input value
dst_offsetconst int32_tinOffset to be applied the output result
activation_minconst int32_tinMinimum value to clamp down the output. Range : int8
activation_maxconst int32_tinMaximum value to clamp up the output. Range : int8
row_address_offsetconst int32_tinAddress offset between rows in output. NOTE: Only used for MVEI extension.
lhs_cols_offsetconst int32_tinColumn offset between subsequent lhs_rows
Returns of arm_nn_mat_mult_nt_t_1x1_out_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

General Matrix-multiplication function with per-channel requantization and int16 input (LHS) and output.

Include/arm_nnsupportfunctions.h:1155

arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s16(
const int16_t *lhs,
const int8_t *rhs,
const cmsis_nn_bias_data *bias_data,
int16_t *dst,
const int32_t *dst_multipliers,
const int32_t *dst_shifts,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t activation_min,
const int32_t activation_max,
const int32_t row_address_offset
)

General Matrix-multiplication function with per-channel requantization and int16 input (LHS) and output. This function assumes:

  • LHS input matrix NOT transposed (nt)
  • RHS input matrix transposed (t)

MVE implementation only.

Parameters of arm_nn_mat_mult_nt_t_s16
NameTypeDirectionDescription
lhsconst int16_t *inPointer to the LHS input matrix
rhsconst int8_t *inPointer to the RHS input matrix
bias_dataconst cmsis_nn_bias_data *inPointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct.
dstint16_t *outPointer to the output matrix with "m" rows and "n" columns
dst_multipliersconst int32_t *inPointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shiftsconst int32_t *inPointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rowsconst int32_tinNumber of LHS input rows
rhs_rowsconst int32_tinNumber of RHS input rows
rhs_colsconst int32_tinNumber of LHS/RHS input columns
activation_minconst int32_tinMinimum value to clamp down the output. Range : int16
activation_maxconst int32_tinMaximum value to clamp up the output. Range : int16
row_address_offsetconst int32_tinAddress offset between rows in output. NOTE: Only used for MVEI extension.
Returns of arm_nn_mat_mult_nt_t_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS` or `ARM_CMSIS_NN_NO_IMPL_ERROR` if not for MVE |---row_address_offset---| |____rhs_rows__________________| | | | | --- | --- | | | | | | | | | | lhs_rows | | | | --- | --- | | _______________ | ______________ |
function

General Matrix-multiplication function with int8 input and int32 output.

Include/arm_nnsupportfunctions.h:1189

arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s8_s32(
const int8_t *lhs,
const int8_t *rhs,
int32_t *dst,
const int32_t lhs_rows,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t lhs_offset,
const int32_t dst_idx_offset
)

General Matrix-multiplication function with int8 input and int32 output. This function assumes:

  • LHS input matrix NOT transposed (nt)
  • RHS input matrix transposed (t)
Parameters of arm_nn_mat_mult_nt_t_s8_s32
NameTypeDirectionDescription
lhsconst int8_t *inPointer to the LHS input matrix
rhsconst int8_t *inPointer to the RHS input matrix
dstint32_t *in, outPointer to the output matrix with "m" rows and "n" columns. Accumulated into, so it must be zeroed by the caller before the call
lhs_rowsconst int32_tinNumber of LHS input rows
rhs_rowsconst int32_tinNumber of LHS input columns/RHS input rows
rhs_colsconst int32_tinNumber of RHS input columns
lhs_offsetconst int32_tinOffset to be applied to the LHS input value
dst_idx_offsetconst int32_tinOffset between subsequent output results
Returns of arm_nn_mat_mult_nt_t_s8_s32
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s4 Vector by Matrix (transposed) multiplication

Include/arm_nnsupportfunctions.h:1218

arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s4(
const int8_t *lhs,
const int8_t *packed_rhs,
const int32_t *bias,
int8_t *dst,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)

s4 Vector by Matrix (transposed) multiplication

Parameters of arm_nn_vec_mat_mult_t_s4
NameTypeDirectionDescription
lhsconst int8_t *inInput left-hand side vector
packed_rhsconst int8_t *inInput right-hand side matrix (transposed)
biasconst int32_t *inInput bias
dstint8_t *outOutput vector
lhs_offsetconst int32_tinOffset to be added to the input values of the left-hand side vector. Range: -127 to 128
dst_offsetconst int32_tinOffset to be added to the output values. Range: -127 to 128
dst_multiplierconst int32_tinOutput multiplier
dst_shiftconst int32_tinOutput shift
rhs_colsconst int32_tinNumber of columns in the right-hand side input matrix
rhs_rowsconst int32_tinNumber of rows in the right-hand side input matrix
activation_minconst int32_tinMinimum value to clamp the output to. Range: int8
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int8
Returns of arm_nn_vec_mat_mult_t_s4
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s8 Vector by Matrix (transposed) multiplication

Include/arm_nnsupportfunctions.h:1256

arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s8(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *kernel_sum,
const int32_t *bias,
int8_t *dst,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max,
const int32_t address_offset,
const int32_t rhs_offset
)

s8 Vector by Matrix (transposed) multiplication

Parameters of arm_nn_vec_mat_mult_t_s8
NameTypeDirectionDescription
lhsconst int8_t *inInput left-hand side vector
rhsconst int8_t *inInput right-hand side matrix (transposed)
kernel_sumconst int32_t *inKernel sums of the kernels (rhs). See arm_vector_sum_s8 for more info.
biasconst int32_t *inInput bias
dstint8_t *outOutput vector
lhs_offsetconst int32_tinOffset to be added to the input values of the left-hand side vector. Range: -127 to 128
dst_offsetconst int32_tinOffset to be added to the output values. Range: -127 to 128
dst_multiplierconst int32_tinOutput multiplier
dst_shiftconst int32_tinOutput shift
rhs_colsconst int32_tinNumber of columns in the right-hand side input matrix
rhs_rowsconst int32_tinNumber of rows in the right-hand side input matrix
activation_minconst int32_tinMinimum value to clamp the output to. Range: int8
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int8
address_offsetconst int32_tinMemory position offset for dst. First output is stored at 'dst', the second at 'dst + address_offset' and so on. Default value is typically 1.
rhs_offsetconst int32_tinOffset to be added to the input values of the right-hand side vector. Range: -127 to 128
Returns of arm_nn_vec_mat_mult_t_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s8 Vector by Matrix (transposed) multiplication using per channel quantization for output

Include/arm_nnsupportfunctions.h:1297

arm_cmsis_nn_status arm_nn_vec_mat_mult_t_per_ch_s8(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *kernel_sum,
const int32_t *bias,
int8_t *dst,
const int32_t lhs_offset,
const int32_t dst_offset,
const int32_t *dst_multiplier,
const int32_t *dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max,
const int32_t address_offset,
const int32_t rhs_offset
)

s8 Vector by Matrix (transposed) multiplication using per channel quantization for output

Parameters of arm_nn_vec_mat_mult_t_per_ch_s8
NameTypeDirectionDescription
lhsconst int8_t *inInput left-hand side vector
rhsconst int8_t *inInput right-hand side matrix (transposed)
kernel_sumconst int32_t *inKernel sums of the kernels (rhs). See arm_vector_sum_s8 for more info.
biasconst int32_t *inInput bias
dstint8_t *outOutput vector
lhs_offsetconst int32_tinOffset to be added to the input values of the left-hand side vector. Range: -127 to 128
dst_offsetconst int32_tinOffset to be added to the output values. Range: -127 to 128
dst_multiplierconst int32_t *inOutput multipliers
dst_shiftconst int32_t *inOutput shifts
rhs_colsconst int32_tinNumber of columns in the right-hand side input matrix
rhs_rowsconst int32_tinNumber of rows in the right-hand side input matrix
activation_minconst int32_tinMinimum value to clamp the output to. Range: int8
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int8
address_offsetconst int32_tinMemory position offset for dst. First output is stored at 'dst', the second at 'dst + address_offset' and so on. Default value is typically 1.
rhs_offsetconst int32_tinOffset to be added to the input values of the right-hand side vector. Range: -127 to 128
Returns of arm_nn_vec_mat_mult_t_per_ch_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s16 Vector by s8 Matrix (transposed) multiplication

Include/arm_nnsupportfunctions.h:1330

arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s16(
const int16_t *lhs,
const int8_t *rhs,
const int64_t *bias,
int16_t *dst,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)

s16 Vector by s8 Matrix (transposed) multiplication

Parameters of arm_nn_vec_mat_mult_t_s16
NameTypeDirectionDescription
lhsconst int16_t *inInput left-hand side vector
rhsconst int8_t *inInput right-hand side matrix (transposed)
biasconst int64_t *inInput bias
dstint16_t *outOutput vector
dst_multiplierconst int32_tinOutput multiplier
dst_shiftconst int32_tinOutput shift
rhs_colsconst int32_tinNumber of columns in the right-hand side input matrix
rhs_rowsconst int32_tinNumber of rows in the right-hand side input matrix
activation_minconst int32_tinMinimum value to clamp the output to. Range: int16
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int16
Returns of arm_nn_vec_mat_mult_t_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s16 vector(lhs) by s8 matrix (transposed) multiplication and per channel quant output

Include/arm_nnsupportfunctions.h:1358

arm_cmsis_nn_status arm_nn_vec_mat_mult_t_per_ch_s16(
const int16_t *lhs,
const int8_t *rhs,
const int64_t *bias,
int16_t *dst,
const int32_t *dst_multiplier,
const int32_t *dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)

s16 vector(lhs) by s8 matrix (transposed) multiplication and per channel quant output

Parameters of arm_nn_vec_mat_mult_t_per_ch_s16
NameTypeDirectionDescription
lhsconst int16_t *inInput left-hand side vector
rhsconst int8_t *inInput right-hand side matrix (transposed)
biasconst int64_t *inInput bias
dstint16_t *outOutput vector
dst_multiplierconst int32_t *inPer channel output multiplier. Length of vector is equal to rhs_rows
dst_shiftconst int32_t *inPer channel output shift. Length of vector is equal to rhs_rows
rhs_colsconst int32_tinNumber of columns in the right-hand side input matrix
rhs_rowsconst int32_tinNumber of rows in the right-hand side input matrix
activation_minconst int32_tinMinimum value to clamp the output to. Range: int16
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int16
Returns of arm_nn_vec_mat_mult_t_per_ch_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s16 Vector by s16 Matrix (transposed) multiplication

Include/arm_nnsupportfunctions.h:1386

arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s16_s16(
const int16_t *lhs,
const int16_t *rhs,
const int64_t *bias,
int16_t *dst,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)

s16 Vector by s16 Matrix (transposed) multiplication

Parameters of arm_nn_vec_mat_mult_t_s16_s16
NameTypeDirectionDescription
lhsconst int16_t *inInput left-hand side vector
rhsconst int16_t *inInput right-hand side matrix (transposed)
biasconst int64_t *inInput bias
dstint16_t *outOutput vector
dst_multiplierconst int32_tinOutput multiplier
dst_shiftconst int32_tinOutput shift
rhs_colsconst int32_tinNumber of columns in the right-hand side input matrix
rhs_rowsconst int32_tinNumber of rows in the right-hand side input matrix
activation_minconst int32_tinMinimum value to clamp the output to. Range: int16
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int16
Returns of arm_nn_vec_mat_mult_t_s16_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s8 Vector by Matrix (transposed) multiplication with s16 output

Include/arm_nnsupportfunctions.h:1417

arm_cmsis_nn_status arm_nn_vec_mat_mult_t_svdf_s8(
const int8_t *lhs,
const int8_t *rhs,
int16_t *dst,
const int32_t lhs_offset,
const int32_t scatter_offset,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t activation_min,
const int32_t activation_max
)

s8 Vector by Matrix (transposed) multiplication with s16 output

Parameters of arm_nn_vec_mat_mult_t_svdf_s8
NameTypeDirectionDescription
lhsconst int8_t *inInput left-hand side vector
rhsconst int8_t *inInput right-hand side matrix (transposed)
dstint16_t *outOutput vector
lhs_offsetconst int32_tinOffset to be added to the input values of the left-hand side vector. Range: -127 to 128
scatter_offsetconst int32_tinAddress offset for dst. First output is stored at 'dst', the second at 'dst + scatter_offset' and so on.
dst_multiplierconst int32_tinOutput multiplier
dst_shiftconst int32_tinOutput shift
rhs_colsconst int32_tinNumber of columns in the right-hand side input matrix
rhs_rowsconst int32_tinNumber of rows in the right-hand side input matrix
activation_minconst int32_tinMinimum value to clamp the output to. Range: int16
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int16
Returns of arm_nn_vec_mat_mult_t_svdf_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Depthwise convolution of transposed rhs matrix with 4 lhs matrices.

Include/arm_nnsupportfunctions.h:1453

arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_padded_s8(
const int8_t *lhs,
const int8_t *rhs,
const int32_t lhs_offset,
const int32_t active_ch,
const int32_t total_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t row_x_col,
const int32_t *const output_bias,
int8_t *out
)

Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in padded cases where the padding is -lhs_offset(Range: int8). Dimensions are the same for lhs and rhs.

Parameters of arm_nn_depthwise_conv_nt_t_padded_s8
NameTypeDirectionDescription
lhsconst int8_t *inInput left-hand side matrix
rhsconst int8_t *inInput right-hand side matrix (transposed)
lhs_offsetconst int32_tinLHS matrix offset(input offset). Range: -127 to 128
active_chconst int32_tinSubset of total_ch processed
total_chconst int32_tinNumber of channels in LHS/RHS
out_shiftconst int32_t *inPer channel output shift. Length of vector is equal to number of channels
out_multconst int32_t *inPer channel output multiplier. Length of vector is equal to number of channels
out_offsetconst int32_tinOffset to be added to the output values. Range: -127 to 128
activation_minconst int32_tinMinimum value to clamp the output to. Range: int8
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int8
row_x_colconst uint16_tin(row_dimension * col_dimension) of LHS/RHS matrix
output_biasconst int32_t *constinPer channel output bias. Length of vector is equal to number of channels
outint8_t *outOutput pointer
Returns of arm_nn_depthwise_conv_nt_t_padded_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` if an implementation is available or `ARM_CMSIS_NN_NO_IMPL_ERROR` otherwise
function

Depthwise convolution of transposed rhs matrix with 4 lhs matrices.

Include/arm_nnsupportfunctions.h:1492

arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_s8(
const int32_t *weight_sum_buf,
const int8_t *lhs,
const int8_t *rhs,
const int32_t lhs_offset,
const int32_t active_ch,
const int32_t total_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t row_x_col,
const int32_t *const output_bias,
int8_t *out
)

Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. Dimensions are the same for lhs and rhs.

Parameters of arm_nn_depthwise_conv_nt_t_s8
NameTypeDirectionDescription
weight_sum_bufconst int32_t *inPointer to the weight sum multiplied by lhs_offset and summed bias buffer
lhsconst int8_t *inInput left-hand side matrix
rhsconst int8_t *inInput right-hand side matrix (transposed)
lhs_offsetconst int32_tinLHS matrix offset(input offset). Range: -127 to 128
active_chconst int32_tinSubset of total_ch processed
total_chconst int32_tinNumber of channels in LHS/RHS
out_shiftconst int32_t *inPer channel output shift. Length of vector is equal to number of channels.
out_multconst int32_t *inPer channel output multiplier. Length of vector is equal to number of channels.
out_offsetconst int32_tinOffset to be added to the output values. Range: -127 to 128
activation_minconst int32_tinMinimum value to clamp the output to. Range: int8
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int8
row_x_colconst uint16_tin(row_dimension * col_dimension) of LHS/RHS matrix
output_biasconst int32_t *constinPer channel output bias. Length of vector is equal to number of channels.
outint8_t *outOutput pointer
Returns of arm_nn_depthwise_conv_nt_t_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` if an implementation is available or `ARM_CMSIS_NN_NO_IMPL_ERROR` otherwise
function

Necessary conditions of the planar rule that are cheap to test inline: at most 32 channels and stride 1.

Include/arm_nnsupportfunctions.h:1517

static int32_t arm_nn_depthwise_conv_s8_planar_candidate(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims
)

Necessary conditions of the planar rule that are cheap to test inline: at most 32 channels and stride 1. A caller can skip arm_nn_depthwise_conv_s8_planar() for layers that fail them without changing which layers it takes.

Parameters of arm_nn_depthwise_conv_s8_planar_candidate
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [1, H, W, C_IN]
Returns of arm_nn_depthwise_conv_s8_planar_candidate
Description
1 when the layer may take the planar path, 0 when it cannot.
function

The gate of armconvolves8smallcin(): upscaledims NULL, input depth 1 to 3 with filter depth equal to it, dilation 1, a kernel of at least 1x1 with kernel width…

Include/arm_nnsupportfunctions.h:1536

static int32_t arm_nn_is_convolve_s8_small_cin(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
const cmsis_nn_dims *upscale_dims
)

The gate of arm_convolve_s8_small_cin(): upscale_dims NULL, input depth 1 to 3 with filter depth equal to it, dilation 1, a kernel of at least 1x1 with kernel width x depth at most 16 and at most 48 values, and a positive multiple of 4 output channels. Plain C; it evaluates the same on every build.

Parameters of arm_nn_is_convolve_s8_small_cin
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, CK]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
upscale_dimsconst cmsis_nn_dims *inUpscale tensor dimensions, or NULL
Returns of arm_nn_is_convolve_s8_small_cin
Description
1 when the layer is in the gate, 0 otherwise.
function

The gate of armconvolves83x3c16s1(): upscaledims NULL, input and filter depth 16, a 3x3 kernel, and stride and dilation 1.

Include/arm_nnsupportfunctions.h:1562

static int32_t arm_nn_is_convolve_s8_3x3_c16_s1(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *upscale_dims
)

The gate of arm_convolve_s8_3x3_c16_s1(): upscale_dims NULL, input and filter depth 16, a 3x3 kernel, and stride and dilation 1. Plain C; it evaluates the same on every build.

Parameters of arm_nn_is_convolve_s8_3x3_c16_s1
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, CK]
upscale_dimsconst cmsis_nn_dims *inUpscale tensor dimensions, or NULL
Returns of arm_nn_is_convolve_s8_3x3_c16_s1
Description
1 when the layer is in the gate, 0 otherwise.
function

The group check of armconvolves8(), for its direct entries: with groups = CIN / filter C, CIN or COUT is not a multiple of groups.

Include/arm_nnsupportfunctions.h:1582

static int32_t arm_nn_convolve_s8_groups_invalid(
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

The group check of arm_convolve_s8(), for its direct entries: with groups = C_IN / filter C, C_IN or C_OUT is not a multiple of groups. A filter C of zero or above C_IN gives no group count and is not reported.

Parameters of arm_nn_convolve_s8_groups_invalid
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, CK]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_nn_convolve_s8_groups_invalid
Description
1 when `arm_convolve_s8()` reports the group count as an argument error, 0 otherwise.
function

Plane size in bytes that armnndepthwiseconvs8planar() needs for a layer, or -1 when the layer is not one it takes.

Include/arm_nnsupportfunctions.h:1601

int32_t arm_nn_depthwise_conv_s8_planar_bytes(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Plane size in bytes that arm_nn_depthwise_conv_s8_planar() needs for a layer, or -1 when the layer is not one it takes. The rule is plain C and evaluates the same on every build.

Parameters of arm_nn_depthwise_conv_s8_planar_bytes
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [1, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_nn_depthwise_conv_s8_planar_bytes
Description
The plane size in bytes, or -1.
function

s8 depthwise convolution with channel multiplier 1 and stride 1, vectorized across the output pixels of one channel plane instead of across channels.

Include/arm_nnsupportfunctions.h:1626

arm_cmsis_nn_status arm_nn_depthwise_conv_s8_planar(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input,
const cmsis_nn_dims *filter_dims,
const int8_t *kernel,
const cmsis_nn_dims *output_dims,
int8_t *output
)

s8 depthwise convolution with channel multiplier 1 and stride 1, vectorized across the output pixels of one channel plane instead of across channels. It serves the few-channel and 1xk layers of arm_depthwise_conv_s8_opt(), with the same scratch buffer and weight sums.

Parameters of arm_nn_depthwise_conv_s8_planar
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outScratch buffer of `arm_depthwise_conv_s8_opt_get_buffer_size()` bytes
weight_sum_ctxconst cmsis_nn_context *inPer-channel weight sums from `arm_depthwise_convolve_weight_sum()`, bias included
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization parameters
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [1, H, W, C_IN]
inputconst int8_t *inInput data pointer
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
kernelconst int8_t *inFilter data pointer
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
outputint8_t *outOutput data pointer
Returns of arm_nn_depthwise_conv_s8_planar
Description
`ARM_CMSIS_NN_SUCCESS` when the layer was computed, or `ARM_CMSIS_NN_NO_IMPL_ERROR` when it is not one this path takes or its plane does not fit in ctx->size (then nothing is written), or MVE is not available.
function

Depthwise convolution of transposed rhs matrix with 4 lhs matrices.

Include/arm_nnsupportfunctions.h:1663

arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_s4(
const int8_t *lhs,
const int8_t *rhs,
const int32_t lhs_offset,
const int32_t active_ch,
const int32_t total_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t row_x_col,
const int32_t *const output_bias,
int8_t *out
)

Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. rhs consists of packed int4 data. Dimensions are the same for lhs and rhs.

Parameters of arm_nn_depthwise_conv_nt_t_s4
NameTypeDirectionDescription
lhsconst int8_t *inInput left-hand side matrix
rhsconst int8_t *inInput right-hand side matrix (transposed). Consists of int4 data packed in an int8 buffer.
lhs_offsetconst int32_tinLHS matrix offset(input offset). Range: -127 to 128
active_chconst int32_tinSubset of total_ch processed
total_chconst int32_tinNumber of channels in LHS/RHS
out_shiftconst int32_t *inPer channel output shift. Length of vector is equal to number of channels.
out_multconst int32_t *inPer channel output multiplier. Length of vector is equal to number of channels.
out_offsetconst int32_tinOffset to be added to the output values. Range: -127 to 128
activation_minconst int32_tinMinimum value to clamp the output to. Range: int8
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int8
row_x_colconst uint16_tin(row_dimension * col_dimension) of LHS/RHS matrix
output_biasconst int32_t *constinPer channel output bias. Length of vector is equal to number of channels.
outint8_t *outOutput pointer
Returns of arm_nn_depthwise_conv_nt_t_s4
Description
The function returns one of the two - Updated output pointer if an implementation is available - NULL if no implementation is available.
function

Depthwise convolution of transposed rhs matrix with 4 lhs matrices.

Include/arm_nnsupportfunctions.h:1699

int16_t * arm_nn_depthwise_conv_nt_t_s16(
const int16_t *lhs,
const int8_t *rhs,
const uint16_t num_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t activation_min,
const int32_t activation_max,
const uint16_t row_x_col,
const int64_t *const output_bias,
int16_t *out
)

Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. Dimensions are the same for lhs and rhs.

Parameters of arm_nn_depthwise_conv_nt_t_s16
NameTypeDirectionDescription
lhsconst int16_t *inInput left-hand side matrix
rhsconst int8_t *inInput right-hand side matrix (transposed)
num_chconst uint16_tinNumber of channels in LHS/RHS
out_shiftconst int32_t *inPer channel output shift. Length of vector is equal to number of channels.
out_multconst int32_t *inPer channel output multiplier. Length of vector is equal to number of channels.
activation_minconst int32_tinMinimum value to clamp the output to. Range: int8
activation_maxconst int32_tinMaximum value to clamp the output to. Range: int8
row_x_colconst uint16_tin(row_dimension * col_dimension) of LHS/RHS matrix
output_biasconst int64_t *constinPer channel output bias. Length of vector is equal to number of channels.
outint16_t *outOutput pointer
Returns of arm_nn_depthwise_conv_nt_t_s16
Description
The function returns one of the two - Updated output pointer if an implementation is available - NULL if no implementation is available.
function

Row of s8 scalars multiplicated with a s8 matrix ad accumulated into a s32 rolling scratch buffer.

Include/arm_nnsupportfunctions.h:1735

arm_cmsis_nn_status arm_nn_transpose_conv_row_s8_s32(
const int8_t *lhs,
const int8_t *rhs,
int32_t *output_start,
const int32_t output_index,
const int32_t output_max,
const int32_t rhs_rows,
const int32_t rhs_cols,
const int32_t input_channels,
const int32_t output_channels,
const int32_t lhs_offset,
const int32_t row_offset,
const int32_t input_x,
const int32_t stride_x,
const int32_t skip_row_top,
const int32_t skip_row_bottom
)

Row of s8 scalars multiplicated with a s8 matrix ad accumulated into a s32 rolling scratch buffer. Helpfunction for transposed convolution.

Parameters of arm_nn_transpose_conv_row_s8_s32
NameTypeDirectionDescription
lhsconst int8_t *inInput left-hand side scalars
rhsconst int8_t *inInput right-hand side matrix
output_startint32_t *outOutput buffer start
output_indexconst int32_tinOutput buffer current index
output_maxconst int32_tinOutput buffer size
rhs_rowsconst int32_tinNumber of rows in rhs matrix
rhs_colsconst int32_tinNumber of columns in rhs matrix
input_channelsconst int32_tinNumber of input channels
output_channelsconst int32_tinNumber of output channels
lhs_offsetconst int32_tinOffset added to lhs before multiplication
row_offsetconst int32_tinAddress offset between each row of data output
input_xconst int32_tinLength of lhs scalar row.
stride_xconst int32_tinAddress offset between each scalar-matrix multiplication result.
skip_row_topconst int32_tinSkip rows on top of the filter, used for padding.
skip_row_bottomconst int32_tinSkip rows in the bottom of the filter, used for padding.
Returns of arm_nn_transpose_conv_row_s8_s32
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

Read 2 s16 elements and post increment pointer.

Include/arm_nnsupportfunctions.h:1756

static int32_t arm_nn_read_q15x2_ia(const int16_t **in_q15)

Read 2 s16 elements and post increment pointer.

Parameters of arm_nn_read_q15x2_ia
NameTypeDirectionDescription
in_q15const int16_t **in, outPointer to pointer that holds address of input. Advanced past the elements read.
Returns of arm_nn_read_q15x2_ia
Description
q31 value
function

Read 4 s8 from s8 pointer and post increment pointer.

Include/arm_nnsupportfunctions.h:1771

static int32_t arm_nn_read_s8x4_ia(const int8_t **in_s8)

Read 4 s8 from s8 pointer and post increment pointer.

Parameters of arm_nn_read_s8x4_ia
NameTypeDirectionDescription
in_s8const int8_t **in, outPointer to pointer that holds address of input. Advanced past the elements read.
Returns of arm_nn_read_s8x4_ia
Description
q31 value
function

Read 2 s8 from s8 pointer and post increment pointer.

Include/arm_nnsupportfunctions.h:1785

static int32_t arm_nn_read_s8x2_ia(const int8_t **in_s8)

Read 2 s8 from s8 pointer and post increment pointer.

Parameters of arm_nn_read_s8x2_ia
NameTypeDirectionDescription
in_s8const int8_t **in, outPointer to pointer that holds address of input. Advanced past the elements read.
Returns of arm_nn_read_s8x2_ia
Description
q31 value
function

Read 2 int16 values from int16 pointer.

Include/arm_nnsupportfunctions.h:1799

static int32_t arm_nn_read_s16x2(const int16_t *in)

Read 2 int16 values from int16 pointer.

Parameters of arm_nn_read_s16x2
NameTypeDirectionDescription
inconst int16_t *inpointer to address of input.
Returns of arm_nn_read_s16x2
Description
s32 value
function

Read 4 s8 values.

Include/arm_nnsupportfunctions.h:1812

static int32_t arm_nn_read_s8x4(const int8_t *in_s8)

Read 4 s8 values.

Parameters of arm_nn_read_s8x4
NameTypeDirectionDescription
in_s8const int8_t *inpointer to address of input.
Returns of arm_nn_read_s8x4
Description
s32 value
function

Read 2 s8 values.

Include/arm_nnsupportfunctions.h:1824

static int32_t arm_nn_read_s8x2(const int8_t *in_s8)

Read 2 s8 values.

Parameters of arm_nn_read_s8x2
NameTypeDirectionDescription
in_s8const int8_t *inpointer to address of input.
Returns of arm_nn_read_s8x2
Description
s32 value
function

Write four s8 to s8 pointer and increment pointer afterwards.

Include/arm_nnsupportfunctions.h:1837

static void arm_nn_write_s8x4_ia(int8_t **in, int32_t value)

Write four s8 to s8 pointer and increment pointer afterwards.

Parameters of arm_nn_write_s8x4_ia
NameTypeDirectionDescription
inint8_t **in, outDouble pointer to destination. Advanced past the bytes written.
valueint32_tinFour bytes to copy
function

memset optimized for MVE

Include/arm_nnsupportfunctions.h:1850

static void arm_memset_s8(int8_t *dst, const int8_t val, uint32_t block_size)

memset optimized for MVE

Parameters of arm_memset_s8
NameTypeDirectionDescription
dstint8_t *in, outDestination pointer
valconst int8_tinValue to set
block_sizeuint32_tinNumber of bytes to copy.
function

memset optimized for MVE for 16-bit data.

Include/arm_nnsupportfunctions.h:1873

static void arm_memset_s16(int16_t *dst, const int16_t val, uint32_t block_size)

memset optimized for MVE for 16-bit data.

Parameters of arm_memset_s16
NameTypeDirectionDescription
dstint16_t *in, outDestination pointer.
valconst int16_tin16-bit value to set.
block_sizeuint32_tinNumber of int16_t values to set.
function

Matrix-multiplication function for convolution with per-channel requantization and 4 bit weights.

Include/arm_nnsupportfunctions.h:2094

int8_t * arm_nn_mat_mult_kernel_s4_s16(
const int8_t *input_a,
const int16_t *input_b,
const uint16_t output_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int32_t activation_min,
const int32_t activation_max,
const int32_t num_col_a,
const int32_t *const output_bias,
int8_t *out_0
)

Matrix-multiplication function for convolution with per-channel requantization and 4 bit weights.

This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.

Parameters of arm_nn_mat_mult_kernel_s4_s16
NameTypeDirectionDescription
input_aconst int8_t *inpointer to operand A, int8 packed with 2x int4.
input_bconst int16_t *inpointer to operand B, always consists of 2 vectors.
output_chconst uint16_tinnumber of rows of A
out_shiftconst int32_t *inpointer to per output channel requantization shift parameter.
out_multconst int32_t *inpointer to per output channel requantization multiplier parameter.
out_offsetconst int32_tinoutput tensor offset.
activation_minconst int32_tinminimum value to clamp the output to. Range : int8
activation_maxconst int32_tinmaximum value to clamp the output to. Range : int8
num_col_aconst int32_tinnumber of columns of A
output_biasconst int32_t *constinper output channel bias. Range : int32
out_0int8_t *in, outpointer to output
Returns of arm_nn_mat_mult_kernel_s4_s16
Description
The function returns one of the two 1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available.
function

Matrix-multiplication function for convolution with per-channel requantization.

Include/arm_nnsupportfunctions.h:2128

int8_t * arm_nn_mat_mult_kernel_s8_s16(
const int8_t *input_a,
const int16_t *input_b,
const uint16_t output_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int16_t activation_min,
const int16_t activation_max,
const int32_t num_col_a,
const int32_t aligned_num_col_a,
const int32_t *const output_bias,
int8_t *out_0
)

Matrix-multiplication function for convolution with per-channel requantization.

This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.

Parameters of arm_nn_mat_mult_kernel_s8_s16
NameTypeDirectionDescription
input_aconst int8_t *inpointer to operand A
input_bconst int16_t *inpointer to operand B, always consists of 2 vectors.
output_chconst uint16_tinnumber of rows of A
out_shiftconst int32_t *inpointer to per output channel requantization shift parameter.
out_multconst int32_t *inpointer to per output channel requantization multiplier parameter.
out_offsetconst int32_tinoutput tensor offset.
activation_minconst int16_tinminimum value to clamp the output to. Range : int8
activation_maxconst int16_tinmaximum value to clamp the output to. Range : int8
num_col_aconst int32_tinnumber of columns of A
aligned_num_col_aconst int32_tinnumber of columns of A aligned by 4
output_biasconst int32_t *constinper output channel bias. Range : int32
out_0int8_t *in, outpointer to output
Returns of arm_nn_mat_mult_kernel_s8_s16
Description
The function returns one of the two 1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available.
function

Matrix-multiplication function for convolution with per-channel requantization, supporting an address offset between rows.

Include/arm_nnsupportfunctions.h:2168

int8_t * arm_nn_mat_mult_kernel_row_offset_s8_s16(
const int8_t *input_a,
const int16_t *input_b,
const uint16_t output_ch,
const int32_t *out_shift,
const int32_t *out_mult,
const int32_t out_offset,
const int16_t activation_min,
const int16_t activation_max,
const int32_t num_col_a,
const int32_t aligned_num_col_a,
const int32_t *const output_bias,
const int32_t row_address_offset,
int8_t *out_0
)

Matrix-multiplication function for convolution with per-channel requantization, supporting an address offset between rows.

This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max.

This function is slighly less performant than arm_nn_mat_mult_kernel_s8_s16, but allows support for grouped convolution. Supported framework: TensorFlow Lite micro.

Parameters of arm_nn_mat_mult_kernel_row_offset_s8_s16
NameTypeDirectionDescription
input_aconst int8_t *inpointer to operand A
input_bconst int16_t *inpointer to operand B, always consists of 2 vectors.
output_chconst uint16_tinnumber of rows of A
out_shiftconst int32_t *inpointer to per output channel requantization shift parameter.
out_multconst int32_t *inpointer to per output channel requantization multiplier parameter.
out_offsetconst int32_tinoutput tensor offset.
activation_minconst int16_tinminimum value to clamp the output to. Range : int8
activation_maxconst int16_tinmaximum value to clamp the output to. Range : int8
num_col_aconst int32_tinnumber of columns of A
aligned_num_col_aconst int32_tinnumber of columns of A aligned by 4
output_biasconst int32_t *constinper output channel bias. Range : int32
row_address_offsetconst int32_tinaddress offset between rows in the output
out_0int8_t *in, outpointer to output
Returns of arm_nn_mat_mult_kernel_row_offset_s8_s16
Description
The function returns one of the two 1. The incremented output pointer for a successful operation or 2. NULL if implementation is not available.
function

Common softmax function for s8 input and s8 or s16 output.

Include/arm_nnsupportfunctions.h:2197

void arm_nn_softmax_common_s8(
const int8_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const int32_t diff_min,
const bool int16_output,
void *output
)

Common softmax function for s8 input and s8 or s16 output.

Parameters of arm_nn_softmax_common_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input tensor
num_rowsconst int32_tinNumber of rows in the input tensor
row_sizeconst int32_tinNumber of elements in each input row
multconst int32_tinInput quantization multiplier
shiftconst int32_tinInput quantization shift within the range [0, 31]
diff_minconst int32_tinMinimum difference with max in row. Used to check if the quantized exponential operation can be performed
int16_outputconst boolinIndicating s8 output if 0 else s16 output
outputvoid *outPointer to the output tensor
function

Saturating doubling high multiply.

Include/arm_nnsupportfunctions.h:2234

static int32_t arm_nn_doubling_high_mult(const int32_t m1, const int32_t m2)

Saturating doubling high multiply. Result matches NEON instruction VQRDMULH.

Parameters of arm_nn_doubling_high_mult
NameTypeDirectionDescription
m1const int32_tinMultiplicand. Range: {NN_Q31_MIN, NN_Q31_MAX}
m2const int32_tinMultiplier. Range: {NN_Q31_MIN, NN_Q31_MAX}
Returns of arm_nn_doubling_high_mult
Description
Result of multiplication.
function

Doubling high multiply without saturation.

Include/arm_nnsupportfunctions.h:2272

static int32_t arm_nn_doubling_high_mult_no_sat(int32_t m1, int32_t m2)

Doubling high multiply without saturation. This is intended for requantization where the scale is a positive integer.

Parameters of arm_nn_doubling_high_mult_no_sat
NameTypeDirectionDescription
m1int32_tinMultiplicand. Range: {NN_Q31_MIN, NN_Q31_MAX}
m2int32_tinMultiplier Range: {NN_Q31_MIN, NN_Q31_MAX}
Returns of arm_nn_doubling_high_mult_no_sat
Description
Result of multiplication.
function

Rounding divide by power of two.

Include/arm_nnsupportfunctions.h:2323

static int32_t arm_nn_divide_by_power_of_two(const int32_t dividend, const int32_t exponent)

Rounding divide by power of two.

Parameters of arm_nn_divide_by_power_of_two
NameTypeDirectionDescription
dividendconst int32_tin- Dividend
exponentconst int32_tin- Divisor = power(2, exponent) Range: [0, 31]
Returns of arm_nn_divide_by_power_of_two
Description
Rounded result of division. Midpoint is rounded away from zero.
function

Rounding divide by power of two for non-negative values.

Include/arm_nnsupportfunctions.h:2378

static int32_t arm_nn_nonneg_divide_by_pot_s32(int32_t dividend, int32_t exponent)

Rounding divide by power of two for non-negative values.

Parameters of arm_nn_nonneg_divide_by_pot_s32
NameTypeDirectionDescription
dividendint32_tin- Dividend (assumed to be non-negative)
exponentint32_tin- Divisor = power(2, exponent) Range: [0, 31]
Returns of arm_nn_nonneg_divide_by_pot_s32
Description
Rounded result of division. Midpoint is rounded away from zero.
function

Requantize a given value.

Include/arm_nnsupportfunctions.h:2416

static int32_t arm_nn_requantize(const int32_t val, const int32_t multiplier, const int32_t shift)

Requantize a given value.

Essentially returns (val * multiplier)/(2 ^ shift) with different rounding depending if CMSIS_NN_USE_SINGLE_ROUNDING is defined or not.

Parameters of arm_nn_requantize
NameTypeDirectionDescription
valconst int32_tinValue to be requantized
multiplierconst int32_tinMultiplier. Range {NN_Q31_MIN + 1, Q32_MAX}
shiftconst int32_tinShift. Range: {-31, 30} Default branch: If shift is positive left shift 'val * multiplier' with shift If shift is negative right shift 'val * multiplier' with abs(shift) Single round branch: Input for total_shift in divide by '2 ^ total_shift'
Returns of arm_nn_requantize
Description
Default branch: Returns (val * multiplier) with rounding divided by (2 ^ shift) with rounding Single round branch: Returns (val * multiplier)/(2 ^ (31 - shift)) with rounding
function

Requantize a given 64 bit value.

Include/arm_nnsupportfunctions.h:2453

static int32_t arm_nn_requantize_s64(const int64_t val, const int32_t reduced_multiplier, const int32_t shift)

Requantize a given 64 bit value.

Parameters of arm_nn_requantize_s64
NameTypeDirectionDescription
valconst int64_tinValue to be requantized in the range {-(1<<47)} to {(1<<47) - 1}
reduced_multiplierconst int32_tinReduced multiplier in the range {NN_Q31_MIN + 1, Q32_MAX} to {Q16_MIN + 1, Q16_MAX}
shiftconst int32_tinLeft or right shift for 'val * multiplier' in the range {-31} to {7}
Returns of arm_nn_requantize_s64
Description
Returns (val * multiplier)/(2 ^ shift)
function

Saturating left shift for int16t.

Include/arm_nnsupportfunctions.h:2471

static int16_t arm_nn_sat_lshift_s16(int16_t x, int shift)

Saturating left shift for int16_t.

Parameters of arm_nn_sat_lshift_s16
NameTypeDirectionDescription
xint16_tinvalue to be shifted
shiftintinNonpositive values return x; positive values multiply by 2^shift with s16 saturation.
Returns of arm_nn_sat_lshift_s16
Description
shifted value
function

Saturating Rounding Doubling High Mul (s16).

Include/arm_nnsupportfunctions.h:2490

static int16_t arm_nn_sqrdmulh_s16(int16_t a, int16_t b)

Saturating Rounding Doubling High Mul (s16).

Matches NEON SQRDMULH s16

Parameters of arm_nn_sqrdmulh_s16
NameTypeDirectionDescription
aint16_tinMultiplicand
bint16_tinMultiplier
Returns of arm_nn_sqrdmulh_s16
Description
Result of multiplication.
function

Saturating Non-rounded Doubling High Mul (s16).

Include/arm_nnsupportfunctions.h:2510

static int16_t arm_nn_sqdmulh_s16(int16_t a, int16_t b)

Saturating Non-rounded Doubling High Mul (s16).

Matches NEON SQDMULH s16

Parameters of arm_nn_sqdmulh_s16
NameTypeDirectionDescription
aint16_tinMultiplicand
bint16_tinMultiplier
Returns of arm_nn_sqdmulh_s16
Description
Result of multiplication.
function

Rounding divide by power of two (s16), midpoint away from zero.

Include/arm_nnsupportfunctions.h:2530

static int16_t arm_nn_divide_by_power_of_two_s16(int16_t x, int exponent)

Rounding divide by power of two (s16), midpoint away from zero.

Mirrors arm_nn_divide_by_power_of_two() semantics for s16.

Parameters of arm_nn_divide_by_power_of_two_s16
NameTypeDirectionDescription
xint16_tinDividend
exponentintinDivisor = power(2, exponent) Range: [0, 15]
Returns of arm_nn_divide_by_power_of_two_s16
Description
Rounded result of division. Midpoint is rounded away from zero.
function

memcpy optimized for MVE

Include/arm_nnsupportfunctions.h:2545

static void arm_memcpy_s8(int8_t *dst, const int8_t *src, uint32_t block_size)

memcpy optimized for MVE

Parameters of arm_memcpy_s8
NameTypeDirectionDescription
dstint8_t *in, outDestination pointer
srcconst int8_t *inSource pointer.
block_sizeuint32_tinNumber of bytes to copy.
function

memcpy optimized for MVE

Include/arm_nnsupportfunctions.h:2569

static void arm_memcpy_s16(int16_t *dst, const int16_t *src, uint32_t block_size)

memcpy optimized for MVE

Parameters of arm_memcpy_s16
NameTypeDirectionDescription
dstint16_t *in, outDestination pointer
srcconst int16_t *inSource pointer.
block_sizeuint32_tinNumber of values to copy.
function

memcpy optimized for MVE

Include/arm_nnsupportfunctions.h:2581

static void arm_memcpy_s32(int32_t *dst, const int32_t *src, uint32_t block_size)

memcpy optimized for MVE

Parameters of arm_memcpy_s32
NameTypeDirectionDescription
dstint32_t *in, outDestination pointer
srcconst int32_t *inSource pointer.
block_sizeuint32_tinNumber of values to copy.
function

memcpy wrapper for int16

Include/arm_nnsupportfunctions.h:2593

static void arm_memcpy_q15(int16_t *dst, const int16_t *src, uint32_t block_size)

memcpy wrapper for int16

Parameters of arm_memcpy_q15
NameTypeDirectionDescription
dstint16_t *in, outDestination pointer
srcconst int16_t *inSource pointer.
block_sizeuint32_tinNumber of bytes to copy.
function

Fixed-point exp() of a non-positive value.

Include/arm_nnsupportfunctions.h:2865

static int32_t arm_nn_exp_on_negative_values(int32_t val)

Fixed-point exp() of a non-positive value.

Parameters of arm_nn_exp_on_negative_values
NameTypeDirectionDescription
valint32_tinInput in Q5.26 fixed point. Must be less than or equal to 0
Returns of arm_nn_exp_on_negative_values
Description
exp(val) in Q0.31 fixed point. Returns NN_Q31_MAX when `val` is 0.
function

Saturating multiply by a power of two.

Include/arm_nnsupportfunctions.h:2905

static int32_t arm_nn_mult_by_power_of_two(const int32_t val, const int32_t exp)

Saturating multiply by a power of two.

Parameters of arm_nn_mult_by_power_of_two
NameTypeDirectionDescription
valconst int32_tinValue to be multiplied
expconst int32_tinExponent. Multiplier = power(2, exp)
Returns of arm_nn_mult_by_power_of_two
Description
val * 2^exp saturated to the int32 range
function

Fixed-point 1 / (1 + x) for x in [0, 1), computed with Newton-Raphson iterations.

Include/arm_nnsupportfunctions.h:2920

static int32_t arm_nn_one_over_one_plus_x_for_x_in_0_1(int32_t val)

Fixed-point 1 / (1 + x) for x in [0, 1), computed with Newton-Raphson iterations.

Parameters of arm_nn_one_over_one_plus_x_for_x_in_0_1
NameTypeDirectionDescription
valint32_tinx in Q0.31 fixed point. Range: [0, NN_Q31_MAX]
Returns of arm_nn_one_over_one_plus_x_for_x_in_0_1
Description
1 / (1 + x) in Q0.31 fixed point
function

Write 2 s16 elements and post increment pointer.

Include/arm_nnsupportfunctions.h:2942

static void arm_nn_write_q15x2_ia(int16_t **dest_q15, int32_t src_q31)

Write 2 s16 elements and post increment pointer.

Parameters of arm_nn_write_q15x2_ia
NameTypeDirectionDescription
dest_q15int16_t **in, outPointer to pointer that holds address of destination. Advanced past the elements written.
src_q31int32_tinInput value to be written.
function

Write 2 s8 elements and post increment pointer.

Include/arm_nnsupportfunctions.h:2955

static void arm_nn_write_s8x2_ia(int8_t **dst, int16_t src)

Write 2 s8 elements and post increment pointer.

Parameters of arm_nn_write_s8x2_ia
NameTypeDirectionDescription
dstint8_t **in, outPointer to pointer that holds address of destination. Advanced past the elements written.
srcint16_tinInput value to be written.
function

Get dimension value at specific index.

Include/arm_nnsupportfunctions.h:2969

static int32_t arm_cmsis_nn_dim_at(const cmsis_nn_dims *dims, int32_t index)

Get dimension value at specific index.

Parameters of arm_cmsis_nn_dim_at
NameTypeDirectionDescription
dimsconst cmsis_nn_dims *inPointer to `cmsis_nn_dims` structure
indexint32_tinIndex of dimension to get
Returns of arm_cmsis_nn_dim_at
Description
Dimension value at specified index
function

Calculate the product of all dimensions in a shape array.

Include/arm_nnsupportfunctions.h:2994

static size_t arm_cmsis_nn_shape_product(const int32_t *shape, int32_t length)

Calculate the product of all dimensions in a shape array.

Parameters of arm_cmsis_nn_shape_product
NameTypeDirectionDescription
shapeconst int32_t *inPointer to array containing shape dimensions
lengthint32_tinNumber of dimensions in the shape array
Returns of arm_cmsis_nn_shape_product
Description
Product of all dimensions
function

Update LSTM function for an iteration step using s8 input and output, and s16 internally.

Include/arm_nnsupportfunctions.h:3022

arm_cmsis_nn_status arm_nn_lstm_step_s8(
const int8_t *data_in,
const int8_t *hidden_in,
int8_t *hidden_out,
const cmsis_nn_lstm_params *params,
cmsis_nn_lstm_context *buffers,
const int32_t batch_offset
)

Update LSTM function for an iteration step using s8 input and output, and s16 internally.

Parameters of arm_nn_lstm_step_s8
NameTypeDirectionDescription
data_inconst int8_t *inData input pointer
hidden_inconst int8_t *inHidden state/ recurrent input pointer
hidden_outint8_t *outHidden state/ recurrent output pointer
paramsconst cmsis_nn_lstm_params *inStruct containg all information about the lstm operator, see arm_nn_types.
bufferscmsis_nn_lstm_context *in, outStruct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types.
batch_offsetconst int32_tinNumber of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps.
Returns of arm_nn_lstm_step_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

Update LSTM function for an iteration step using s16 input and output, and s16 internally.

Include/arm_nnsupportfunctions.h:3046

arm_cmsis_nn_status arm_nn_lstm_step_s16(
const int16_t *data_in,
const int16_t *hidden_in,
int16_t *hidden_out,
const cmsis_nn_lstm_params *params,
cmsis_nn_lstm_context *buffers,
const int32_t batch_offset
)

Update LSTM function for an iteration step using s16 input and output, and s16 internally.

Parameters of arm_nn_lstm_step_s16
NameTypeDirectionDescription
data_inconst int16_t *inData input pointer
hidden_inconst int16_t *inHidden state/ recurrent input pointer
hidden_outint16_t *outHidden state/ recurrent output pointer
paramsconst cmsis_nn_lstm_params *inStruct containg all information about the lstm operator, see arm_nn_types.
bufferscmsis_nn_lstm_context *in, outStruct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types.
batch_offsetconst int32_tinNumber of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps.
Returns of arm_nn_lstm_step_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

Updates a LSTM gate for an iteration step of LSTM function, int8x816 version.

Include/arm_nnsupportfunctions.h:3067

arm_cmsis_nn_status arm_nn_lstm_calculate_gate_s8_s16(
const int8_t *data_in,
const int8_t *hidden_in,
const cmsis_nn_lstm_gate *gate_data,
const cmsis_nn_lstm_params *params,
int16_t *output,
const int32_t batch_offset
)

Updates a LSTM gate for an iteration step of LSTM function, int8x8_16 version.

Parameters of arm_nn_lstm_calculate_gate_s8_s16
NameTypeDirectionDescription
data_inconst int8_t *inData input pointer
hidden_inconst int8_t *inHidden state/ recurrent input pointer
gate_dataconst cmsis_nn_lstm_gate *inStruct containing all information about the gate caluclation, see arm_nn_types.
paramsconst cmsis_nn_lstm_params *inStruct containing all information about the lstm_operation, see arm_nn_types
outputint16_t *outHidden state/ recurrent output pointer
batch_offsetconst int32_tinNumber of timesteps between consecutive batches, see arm_nn_lstm_step_s8.
Returns of arm_nn_lstm_calculate_gate_s8_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

Updates a LSTM gate for an iteration step of LSTM function, int16x816 version.

Include/arm_nnsupportfunctions.h:3088

arm_cmsis_nn_status arm_nn_lstm_calculate_gate_s16(
const int16_t *data_in,
const int16_t *hidden_in,
const cmsis_nn_lstm_gate *gate_data,
const cmsis_nn_lstm_params *params,
int16_t *output,
const int32_t batch_offset
)

Updates a LSTM gate for an iteration step of LSTM function, int16x8_16 version.

Parameters of arm_nn_lstm_calculate_gate_s16
NameTypeDirectionDescription
data_inconst int16_t *inData input pointer
hidden_inconst int16_t *inHidden state/ recurrent input pointer
gate_dataconst cmsis_nn_lstm_gate *inStruct containing all information about the gate caluclation, see arm_nn_types.
paramsconst cmsis_nn_lstm_params *inStruct containing all information about the lstm_operation, see arm_nn_types
outputint16_t *outHidden state/ recurrent output pointer
batch_offsetconst int32_tinNumber of timesteps between consecutive batches, see arm_nn_lstm_step_s16.
Returns of arm_nn_lstm_calculate_gate_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

The result of the multiplication is accumulated to the passed result buffer.

Include/arm_nnsupportfunctions.h:3114

arm_cmsis_nn_status arm_nn_vec_mat_mul_result_acc_s8_s16(
const int8_t *lhs,
const int8_t *rhs,
const int32_t *effective_bias,
int16_t *dst,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t batches,
const int32_t batch_offset
)

The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a “batched” vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).

Parameters of arm_nn_vec_mat_mul_result_acc_s8_s16
NameTypeDirectionDescription
lhsconst int8_t *inBatched vector
rhsconst int8_t *inWeights - input matrix (H(Rows)xW(Columns))
effective_biasconst int32_t *inBias + lhs_offset * kernel_sum term precalculated into a constant vector.
dstint16_t *outOutput
dst_multiplierconst int32_tinMultiplier for quantization
dst_shiftconst int32_tinShift for quantization
rhs_colsconst int32_tinVector/matarix column length
rhs_rowsconst int32_tinRow count of matrix
batchesconst int32_tinBatch size
batch_offsetconst int32_tinNumber of timesteps between consecutive batches in input, see arm_nn_lstm_step_s8. Note that the output is always stored with sequential batches.
Returns of arm_nn_vec_mat_mul_result_acc_s8_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

The result of the multiplication is accumulated to the passed result buffer.

Include/arm_nnsupportfunctions.h:3144

arm_cmsis_nn_status arm_nn_vec_mat_mul_result_acc_s16(
const int16_t *lhs,
const int8_t *rhs,
const int64_t *effective_bias,
int16_t *dst,
const int32_t dst_multiplier,
const int32_t dst_shift,
const int32_t rhs_cols,
const int32_t rhs_rows,
const int32_t batches,
const int32_t batch_offset
)

The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a “batched” vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).

Parameters of arm_nn_vec_mat_mul_result_acc_s16
NameTypeDirectionDescription
lhsconst int16_t *inBatched vector
rhsconst int8_t *inWeights - input matrix (H(Rows)xW(Columns))
effective_biasconst int64_t *inBias + lhs_offset * kernel_sum term precalculated into a constant vector.
dstint16_t *outOutput
dst_multiplierconst int32_tinMultiplier for quantization
dst_shiftconst int32_tinShift for quantization
rhs_colsconst int32_tinVector/matarix column length
rhs_rowsconst int32_tinRow count of matrix
batchesconst int32_tinBatch size
batch_offsetconst int32_tinNumber of timesteps between consecutive batches in input, see arm_nn_lstm_step_s16. Note that the output is always stored with sequential batches.
Returns of arm_nn_vec_mat_mul_result_acc_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s16 elementwise multiplication with s8 output

Include/arm_nnsupportfunctions.h:3171

arm_cmsis_nn_status arm_elementwise_mul_s16_s8(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t block_size,
const int32_t batch_size,
const int32_t batch_offset
)

s16 elementwise multiplication with s8 output

Supported framework: TensorFlow Lite micro

Parameters of arm_elementwise_mul_s16_s8
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input vector 1
input_2_vectconst int16_t *inpointer to input vector 2
outputint8_t *in, outpointer to output vector
out_offsetconst int32_tinoutput offset
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
block_sizeconst int32_tinnumber of samples per batch
batch_sizeconst int32_tinnumber of samples per batch
batch_offsetconst int32_tinNumber of timesteps between consecutive batches in output, see arm_nn_lstm_step_s8. Note that it is assumed that the input is stored with sequential batches.
Returns of arm_elementwise_mul_s16_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise multiplication with s16 output

Include/arm_nnsupportfunctions.h:3197

arm_cmsis_nn_status arm_elementwise_mul_s16_batch_offset(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t block_size,
const int32_t batch_size,
const int32_t batch_offset
)

s16 elementwise multiplication with s16 output

Supported framework: TensorFlow Lite micro

Parameters of arm_elementwise_mul_s16_batch_offset
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input vector 1
input_2_vectconst int16_t *inpointer to input vector 2
outputint16_t *in, outpointer to output vector
out_offsetconst int32_tinoutput offset
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
block_sizeconst int32_tinnumber of samples per batch
batch_sizeconst int32_tinnumber of samples per batch
batch_offsetconst int32_tinNumber of timesteps between consecutive batches in output, see arm_nn_lstm_step_s16. Note that it is assumed that the input is stored with sequential batches.
Returns of arm_elementwise_mul_s16_batch_offset
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise multiplication.

Include/arm_nnsupportfunctions.h:3224

arm_cmsis_nn_status arm_elementwise_mul_acc_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise multiplication. The result of the multiplication is accumulated to the passed result buffer.

Supported framework: TensorFlow Lite micro

Parameters of arm_elementwise_mul_acc_s16
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input vector 1
input_2_vectconst int16_t *inpointer to input vector 2
input_1_offsetconst int32_tinoffset for input 1. Not used.
input_2_offsetconst int32_tinoffset for input 2. Not used.
outputint16_t *in, outpointer to output vector
out_offsetconst int32_tinoutput offset. Not used.
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_elementwise_mul_acc_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

Check if a broadcast is required between 2 cmsisnndims.

Include/arm_nnsupportfunctions.h:3245

static int32_t arm_check_broadcast_required(const cmsis_nn_dims *shape_1, const cmsis_nn_dims *shape_2)

Check if a broadcast is required between 2 cmsis_nn_dims.

Compares each dimension and returns 1 if any dimension does not match. This function does not check that broadcast rules are met.

Parameters of arm_check_broadcast_required
NameTypeDirectionDescription
shape_1const cmsis_nn_dims *inpointer to input tensor 1
shape_2const cmsis_nn_dims *inpointer to input tensor 2
Returns of arm_check_broadcast_required
Description
The function returns 1 if a broadcast is required, or 0 if not.
function

Reports whether the reduced axes of a 4-D tensor form one contiguous block followed by kept axes, as in a NHWC mean over H and W, and gives the flattened sizes.

Include/arm_nnsupportfunctions.h:3310

static int32_t arm_reduce_get_middle_block_from_arrays(
const int32_t in_dims,
const int32_t axis_arr,
int32_t *outer,
int32_t *reduce,
int32_t *inner
)

Reports whether the reduced axes of a 4-D tensor form one contiguous block followed by kept axes, as in a NHWC mean over H and W, and gives the flattened sizes. Axes of size 1 are ignored.

Parameters of arm_reduce_get_middle_block_from_arrays
NameTypeDirectionDescription
in_dimsconst int32_tin4-element array {n, h, w, c}
axis_arrconst int32_tin4-element mask {axis_n, axis_h, axis_w, axis_c}
outerint32_t *outProduct of the dims before the reduced block
reduceint32_t *outProduct of the reduced dims
innerint32_t *outProduct of the dims after the reduced block
Returns of arm_reduce_get_middle_block_from_arrays
Description
1 if the input is [outer, reduce, inner] with the middle dim reduced and inner > 1, otherwise 0
function

One element of armsqrts16tablefree(): the float32 chain the MVE path evaluates per lane, so the two agree bit for bit on any IEEE-754 float32 implementation wi…

Include/arm_nnsupportfunctions.h:3381

static int16_t arm_nn_sqrt_s16_tablefree_element(const int32_t value, const float scale)

One element of arm_sqrt_s16_tablefree(): the float32 chain the MVE path evaluates per lane, so the two agree bit for bit on any IEEE-754 float32 implementation with round-to-nearest-even and a fused multiply-add (fmaf). Every product after the pre-scale either has two uses or feeds an fmaf or a conversion, never another lone multiply, so a compiler allowed to reassociate (-ffast-math) still has no chain to reorder, and no product feeds a bare add, so there is nothing to contract.

Parameters of arm_nn_sqrt_s16_tablefree_element
NameTypeDirectionDescription
valueconst int32_tininput code; values <= 0 give 0
scaleconst floatininput_scale / (output_scale * output_scale) as float32
Returns of arm_nn_sqrt_s16_tablefree_element
Description
trunc(sqrt(value * scale)) saturated to 32767