arm_nn_dw_conv_opt_dilation_supportedfunctionCheck if the dilation, stride and padding of a depthwise layer allow the armdepthwiseconvs8opt() or armdepthwiseconvfasts16() route.
arm_q7_to_q15_with_offsetfunctionConverts the elements from a s8 vector to a s16 vector with an added offset.
arm_nn_depthwise_conv_s8_corefunctionDepthwise conv on an im2col buffer where the input channel equals output channel.
arm_nn_mat_mult_s8functionGeneral Matrix-multiplication function with per-channel requantization.
arm_nn_mat_mult_kernel_s16functionMatrix-multiplication function for convolution with per-channel requantization for 16 bits convolution.
arm_nn_mat_mul_core_1x_s8functionGeneral Vector by Matrix multiplication with requantization and storage of result.
arm_nn_mat_mul_core_1x_s4functionGeneral Vector by Matrix multiplication with requantization, storage of result and int4 weights packed into an int8 buffer.
arm_nn_mat_mul_core_4x_s8functionMatrix-multiplication with requantization & activation function for four rows and one column.
arm_nn_mat_mult_nt_t_s4functionGeneral Matrix-multiplication function with per-channel requantization.
arm_nn_is_convolve_s8_small_cinfunctionThe gate of armconvolves8smallcin(): upscaledims NULL, input depth 1 to 3 with filter depth equal to it, dilation 1, a kernel of at least 1x1 with kernel width…
arm_nn_is_convolve_s8_3x3_c16_s1functionThe gate of armconvolves83x3c16s1(): upscaledims NULL, input and filter depth 16, a 3x3 kernel, and stride and dilation 1.
arm_nn_convolve_s8_groups_invalidfunctionThe group check of armconvolves8(), for its direct entries: with groups = CIN / filter C, CIN or COUT is not a multiple of groups.
arm_nn_depthwise_conv_s8_planar_bytesfunctionPlane size in bytes that armnndepthwiseconvs8planar() needs for a layer, or -1 when the layer is not one it takes.
arm_nn_depthwise_conv_s8_planarfunctions8 depthwise convolution with channel multiplier 1 and stride 1, vectorized across the output pixels of one channel plane instead of across channels.
arm_memset_s16functionmemset optimized for MVE for 16-bit data.
arm_nn_mat_mult_kernel_s4_s16functionMatrix-multiplication function for convolution with per-channel requantization and 4 bit weights.
arm_nn_mat_mult_kernel_s8_s16functionMatrix-multiplication function for convolution with per-channel requantization.
arm_nn_mat_mult_kernel_row_offset_s8_s16functionMatrix-multiplication function for convolution with per-channel requantization, supporting an address offset between rows.
arm_reduce_get_middle_block_from_arraysfunctionReports whether the reduced axes of a 4-D tensor form one contiguous block followed by kept axes, as in a NHWC mean over H and W, and gives the flattened sizes.
arm_nn_sqrt_s16_tablefree_elementfunctionOne element of armsqrts16tablefree(): the float32 chain the MVE path evaluates per lane, so the two agree bit for bit on any IEEE-754 float32 implementation wi…
With ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 minNum via VMINNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MIN, an ordered compare, so its NaN handling depends on operand order: a NaN b is returned, a NaN a is not. Do not rely on NaN suppression on non-MVE builds.
With ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 maxNum via VMAXNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MAX, an ordered compare, so its NaN handling depends on operand order: a NaN b is returned, a NaN a is not. Do not rely on NaN suppression on non-MVE builds.
Both the NaN test and the select are performed on the bit patterns: the test is (bits & 0x7FFF) > 0x7C00 (all-ones exponent, non-zero mantissa), which is integer arithmetic that -ffinite-math-only (implied by the shipped -Ofast) has no license to fold, unlike the former floating-point self-compare x != x (#333 / #334); the bit-pattern select neither expands to an HFmode conditional move (PR target/118460) nor quiets/retags the NaN payload. This helper backs the f16 elementwise clamp and, via arm_nn_clamp_scalar_f16 / arm_nn_clamp_propagate_nan_f16h, the other f16 scalar clamp users arm_svdf_f16, arm_max_pool_f16 / arm_avg_pool_f16, the packed f16 matmul (arm_nn_mat_mult_nt_n_packed_f16), the scalar f16 RELU/RELU6/LEAKY_RELU activation legs, and arm_nn_vector_clamp_f16’s scalar leg (conv/depthwise/transpose-conv f16, the 3x3 depthwise, and arm_nn_maxpool1d_f16) so wherever that scalar clamp runs, a NaN passes through it at every optimization level on the gated toolchains. That is a guarantee about the clamp, not the whole kernel: which builds run the scalar clamp, and whether a NaN survives the rest of the kernel to reach it, is per kernel several of these callers clamp with vmaxnmq/vminnmq on MVE builds (a NaN resolves to a bound there), and arm_max_pool_f16’s max reduction drops a NaN before the clamp. The kernels with a NaN
Parameters
Parameters of arm_nn_propagate_nan_f16h
Name
Type
Direction
Description
x
_Float16
in
Value whose NaN-ness selects the result. Returned unchanged when it is NaN.
Drop-in equivalent of ARM_NN_CLAMP(x, h, l) for scalar _Float16 operands.
Includes the macro’s NaN behaviour: ARM_NN_MIN(NaN, h) is h, so a NaN input resolves to the high bound, exactly as the macro does. Use arm_nn_clamp_propagate_nan_f16h() where TFLite NaN propagation is required.
Scalar f16 clamp with TFLite NaN semantics: NaN passes through unchanged.
Mirrors the MVE idiom in arm_nn_clamp_propagate_nan_mve_f16() (lower bound first, then upper bound, then restore NaN lanes). The NaN restore in arm_nn_propagate_nan_f16h() tests the integer bit pattern, so it holds at every optimization level including the shipped -Ofast; see #333 / #334. Bounds are assumed ordered (l <= h); inverted bounds are unspecified.
Parameters
Parameters of arm_nn_clamp_propagate_nan_f16h
Name
Type
Direction
Description
x
_Float16
in
Value to clamp
l
_Float16
in
Lower bound
h
_Float16
in
Upper bound
Returns
Returns of arm_nn_clamp_propagate_nan_f16h
Description
`x` clamped to [`l`, `h`], or `x` itself when it is NaN
Fold one dimension into a running buffer-size product, reporting overflow as -1.
Buffer-size queries return an int32_t byte count, so the product of the dimensions they multiply has to be rejected as soon as it cannot fit. Folding one factor at a time keeps the accumulator bounded: an accumulator already known to be <= INT32_MAX times a factor <= INT32_MAX cannot exceed about 2^62, so the int64_t accumulator itself never wraps. Chaining raw (int64_t) casts across three or more int32_t dims does not have that property - 65536 * 65536 * 65536 * 65536 is exactly 2^64 and folds back to 0, which would sail through a trailing “> INT32_MAX” test.
Parameters
Parameters of arm_nn_size_mul
Name
Type
Direction
Description
acc
const int64_t
in
Running product, or -1 if an earlier fold already overflowed.
factor
const int64_t
in
Next factor to fold in.
Returns
Returns of arm_nn_size_mul
Description
acc * factor, or -1 if acc is already -1, factor is negative or out of int32_t range, or the product exceeds INT32_MAX.
Byte lanes are masked and shifted in uint32_t so a negative value never feeds a signed left shift (UB); masking before the shift keeps the same bits the old shift-then-mask form kept. Bit-identical for every input. Deliberate divergence from upstream ARM-software/CMSIS-NN, which still carries the signed-shift form do not paste the upstream text back on a sync (issue #357).
Same treatment: the high half is shifted in uint32_t, not int32_t, so a negative v1 is defined; the low half keeps its mask. Bit-identical for every input. Same deliberate upstream divergence as PACK_S8x4_32x1 above.
Check that arm_convolve_1_x_n_s4() handles the horizontal padding of a 1xN convolution.
The kernel places pad.w columns on the left and pad.w + (total_pad % 2) on the right, where total_pad = (output W - 1) * stride.w + filter W - input W, and needs the output columns that read padding to fit in output W. Its padded-column code also assumes that each such column reads at least one input column and that the filter is no wider than the input; otherwise it forms input and filter addresses outside the tensors. On MVE builds another pad placement, or too many padded columns, returns ARM_CMSIS_NN_FAILURE. A VALID layer whose stride leaves trailing input unused (negative total_pad) is therefore rejected, and the wrapper routes it to another convolution. The kernel also computes a single output row, so vertical padding or an output height other than 1 is rejected. A non-positive stride.w is left to the kernel’s argument checks.
Parameters
Parameters of arm_nn_convolve_1_x_n_padding_supported
Name
Type
Direction
Description
conv_params
const cmsis_nn_conv_params *
in
Convolution parameters
input_dims
const cmsis_nn_dims *
in
Input dimensions
filter_dims
const cmsis_nn_dims *
in
Filter dimensions
output_dims
const cmsis_nn_dims *
in
Output dimensions
Returns
Returns of arm_nn_convolve_1_x_n_padding_supported
Description
true when `arm_convolve_1_x_n_s4()` handles the padding, false otherwise.
Check that arm_convolve_1_x_n_s8() accepts the padding and output shape of a 1xN convolution.
The kernel computes a single output row for any pad.w >= 0 and any output width, including an odd total padding, a filter wider than the input and a VALID layer whose stride leaves trailing input unused. It rejects vertical padding, an output height other than 1, a negative pad.w and an empty filter; the wrapper routes those layers to another convolution. A non-positive stride.w is left to the kernel’s argument checks.
Parameters
Parameters of arm_nn_convolve_1_x_n_s8_padding_supported
Name
Type
Direction
Description
conv_params
const cmsis_nn_conv_params *
in
Convolution parameters
filter_dims
const cmsis_nn_dims *
in
Filter dimensions
output_dims
const cmsis_nn_dims *
in
Output dimensions
Returns
Returns of arm_nn_convolve_1_x_n_s8_padding_supported
Description
true when `arm_convolve_1_x_n_s8()` computes the layer, false otherwise.
Count the output columns of a 1xN convolution whose window reads padding.
Output column j reads input columns j * stride.w - pad.w to j * stride.w - pad.w + filter W - 1. The leading columns whose window starts before the input are left-padded; of the others, the trailing columns whose window ends past the input are right-padded. A window can do both only when it is left-padded.
Parameters
Parameters of arm_nn_convolve_1_x_n_padded_columns
Name
Type
Direction
Description
conv_params
const cmsis_nn_conv_params *
in
Convolution parameters. stride.w >= 1 and pad.w >= 0.
input_dims
const cmsis_nn_dims *
in
Input dimensions. w >= 0.
filter_dims
const cmsis_nn_dims *
in
Filter dimensions. w >= 1.
output_dims
const cmsis_nn_dims *
in
Output dimensions. w >= 0.
left_num
int64_t *
out
Number of left-padded output columns, output W at most.
right_num
int64_t *
out
Number of right-padded output columns, output W - left_num at most.
Check if the dilation, stride and padding of a depthwise layer allow the arm_depthwise_conv_s8_opt() or arm_depthwise_conv_fast_s16() route.
Parameters
Parameters of arm_nn_dw_conv_opt_dilation_supported
Name
Type
Direction
Description
dw_conv_params
const cmsis_nn_dw_conv_params *
in
Depthwise convolution parameters
input_dims
const cmsis_nn_dims *
in
Input dimensions
filter_dims
const cmsis_nn_dims *
in
Filter dimensions
output_dims
const cmsis_nn_dims *
in
Output dimensions
Returns
Returns of arm_nn_dw_conv_opt_dilation_supported
Description
true for an undilated layer (dilation 1 in both dimensions), or for a 1D layer dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, no vertical padding, dilation.h == 1 and dilation.w >= 1. false otherwise.
Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with MVE extension.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.
Parameters
Parameters of arm_depthwise_conv_s8_opt_get_buffer_size_mve
Name
Type
Direction
Description
input_dims
const cmsis_nn_dims *
in
Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.
Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with DSP extension.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.
Parameters
Parameters of arm_depthwise_conv_s8_opt_get_buffer_size_dsp
Name
Type
Direction
Description
input_dims
const cmsis_nn_dims *
in
Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.
Matrix-multiplication function for convolution with per-channel requantization for 16 bits convolution.
This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.
Parameters
Parameters of arm_nn_mat_mult_kernel_s16
Name
Type
Direction
Description
input_a
const int8_t *
in
pointer to operand A
input_b
const int16_t *
in
pointer to operand B, always consists of 2 vectors.
output_ch
const int32_t
in
number of rows of A
out_shift
const int32_t *
in
pointer to per output channel requantization shift parameter.
out_mult
const int32_t *
in
pointer to per output channel requantization multiplier parameter.
activation_min
const int32_t
in
minimum value to clamp the output to. Range : int16
activation_max
const int32_t
in
maximum value to clamp the output to. Range : int16
num_col_a
const int32_t
in
number of columns of A
bias_data
const cmsis_nn_bias_data *const
in
pointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct.
out_0
int16_t *
in, out
pointer to output
row_address_offset
const int32_t
in
Address offset between rows in output.
Returns
Returns of arm_nn_mat_mult_kernel_s16
Description
The function returns one of the two
1. The incremented output pointer for a successful operation or
2. NULL if implementation is not available.
General Vector by Matrix multiplication with requantization and storage of result.
Pseudo-code *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in ‘output’
Parameters
Parameters of arm_nn_mat_mul_core_1x_s8
Name
Type
Direction
Description
row_elements
int32_t
in
number of row elements
skipped_row_elements
const int32_t
in
number of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch
row_base_ref
const int8_t *
in
pointer to row operand
col_base_ref
const int8_t *
in
pointer to col operand
out_ch
const int32_t
in
Number of output channels
conv_params
const cmsis_nn_conv_params *
in
Pointer to convolution parameters like offsets and activation values
quant_params
const cmsis_nn_per_channel_quant_params *
in
Pointer to per-channel quantization parameters
bias
const int32_t *
in
Pointer to optional per-channel bias
output
int8_t *
out
Pointer to output where int8 results are stored.
Returns
Returns of arm_nn_mat_mul_core_1x_s8
Description
The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory.
General Vector by Matrix multiplication with requantization, storage of result and int4 weights packed into an int8 buffer.
Pseudo-code as int8 example. Int4 filter data will be unpacked. *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in ‘output’
Parameters
Parameters of arm_nn_mat_mul_core_1x_s4
Name
Type
Direction
Description
row_elements
int32_t
in
number of row elements
skipped_row_elements
const int32_t
in
number of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch
row_base_ref
const int8_t *
in
pointer to row operand
col_base_ref
const int8_t *
in
pointer to col operand as packed int4
out_ch
const int32_t
in
Number of output channels
conv_params
const cmsis_nn_conv_params *
in
Pointer to convolution parameters like offsets and activation values
quant_params
const cmsis_nn_per_channel_quant_params *
in
Pointer to per-channel quantization parameters
bias
const int32_t *
in
Pointer to optional per-channel bias
output
int8_t *
out
Pointer to output where int8 results are stored.
Returns
Returns of arm_nn_mat_mul_core_1x_s4
Description
The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory.
General Matrix-multiplication function with per-channel requantization. This function assumes:
LHS input matrix NOT transposed (nt)
RHS input matrix transposed (t)
RHS is int8 packed with 2x int4
LHS is int8
Parameters
Parameters of arm_nn_mat_mult_nt_t_s4
Name
Type
Direction
Description
lhs
const int8_t *
in
Pointer to the LHS input matrix
rhs
const int8_t *
in
Pointer to the RHS input matrix
bias
const int32_t *
in
Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)
dst
int8_t *
out
Pointer to the output matrix with "m" rows and "n" columns
dst_multipliers
const int32_t *
in
Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shifts
const int32_t *
in
Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rows
const int32_t
in
Number of LHS input rows
rhs_rows
const int32_t
in
Number of RHS input rows
rhs_cols
const int32_t
in
Number of LHS/RHS input columns
lhs_offset
const int32_t
in
Offset to be applied to the LHS input value
dst_offset
const int32_t
in
Offset to be applied the output result
activation_min
const int32_t
in
Minimum value to clamp down the output. Range : int8
activation_max
const int32_t
in
Maximum value to clamp up the output. Range : int8
General Matrix-multiplication function with per-channel requantization. This function assumes:
LHS input matrix NOT transposed (nt)
RHS input matrix transposed (t)
RHS is int8 packed with 2x int4
LHS is int8
LHS/RHS input columns must be even numbered
LHS must be interleaved. Compare to arm_nn_mat_mult_nt_t_s4 where LHS is not interleaved.
Parameters
Parameters of arm_nn_mat_mult_nt_interleaved_t_even_s4
Name
Type
Direction
Description
lhs
const int8_t *
in
Pointer to the LHS input matrix
rhs
const int8_t *
in
Pointer to the RHS input matrix
bias
const int32_t *
in
Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)
dst
int8_t *
out
Pointer to the output matrix with "m" rows and "n" columns
dst_multipliers
const int32_t *
in
Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shifts
const int32_t *
in
Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rows
const int32_t
in
Number of LHS input rows
rhs_rows
const int32_t
in
Number of RHS input rows
rhs_cols
const int32_t
in
Number of LHS/RHS input columns. Note this must be even.
lhs_offset
const int32_t
in
Offset to be applied to the LHS input value
dst_offset
const int32_t
in
Offset to be applied the output result
activation_min
const int32_t
in
Minimum value to clamp down the output. Range : int8
activation_max
const int32_t
in
Maximum value to clamp up the output. Range : int8
lhs_cols_offset
const int32_t
in
Column offset between subsequent lhs_rows
Returns
Returns of arm_nn_mat_mult_nt_interleaved_t_even_s4
General Matrix-multiplication function with per-channel requantization. This function assumes:
LHS input matrix NOT transposed (nt)
RHS input matrix transposed (t)
Parameters
Parameters of arm_nn_mat_mult_nt_t_s8
Name
Type
Direction
Description
weight_sum_buf
const int32_t *
in
Pointer to the weight sum multiplied by lhs_offset and summed bias buffer
lhs
const int8_t *
in
Pointer to the LHS input matrix
rhs
const int8_t *
in
Pointer to the RHS input matrix
bias
const int32_t *
in
Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)
dst
int8_t *
out
Pointer to the output matrix with "m" rows and "n" columns
dst_multipliers
const int32_t *
in
Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shifts
const int32_t *
in
Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rows
const int32_t
in
Number of LHS input rows
rhs_rows
const int32_t
in
Number of RHS input rows
rhs_cols
const int32_t
in
Number of LHS/RHS input columns
lhs_offset
const int32_t
in
Offset to be applied to the LHS input value
dst_offset
const int32_t
in
Offset to be applied the output result
activation_min
const int32_t
in
Minimum value to clamp down the output. Range : int8
activation_max
const int32_t
in
Maximum value to clamp up the output. Range : int8
row_address_offset
const int32_t
in
Address offset between rows in output. NOTE: Only used for MVEI extension.
General Matrix-multiplication function with per-channel requantization. Output is calculated with multiple channels in parallel, rather than multiple output indices in a single channel This function assumes:
LHS input matrix NOT transposed (nt)
RHS input matrix transposed (t)
Parameters
Parameters of arm_nn_mat_mult_nt_t_1x1_out_s8
Name
Type
Direction
Description
weight_sum_buf
const int32_t *
in
Pointer to the weight sum multiplied by lhs_offset and summed bias buffer
lhs
const int8_t *
in
Pointer to the LHS input matrix
rhs
const int8_t *
in
Pointer to the RHS input matrix
bias
const int32_t *
in
Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)
dst
int8_t *
out
Pointer to the output matrix with "m" rows and "n" columns
dst_multipliers
const int32_t *
in
Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shifts
const int32_t *
in
Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rows
const int32_t
in
Number of LHS input rows
rhs_rows
const int32_t
in
Number of RHS input rows
rhs_cols
const int32_t
in
Number of LHS/RHS input columns
lhs_offset
const int32_t
in
Offset to be applied to the LHS input value
dst_offset
const int32_t
in
Offset to be applied the output result
activation_min
const int32_t
in
Minimum value to clamp down the output. Range : int8
activation_max
const int32_t
in
Maximum value to clamp up the output. Range : int8
row_address_offset
const int32_t
in
Address offset between rows in output. NOTE: Only used for MVEI extension.
General Matrix-multiplication function with per-channel requantization and int16 input (LHS) and output. This function assumes:
LHS input matrix NOT transposed (nt)
RHS input matrix transposed (t)
MVE implementation only.
Parameters
Parameters of arm_nn_mat_mult_nt_t_s16
Name
Type
Direction
Description
lhs
const int16_t *
in
Pointer to the LHS input matrix
rhs
const int8_t *
in
Pointer to the RHS input matrix
bias_data
const cmsis_nn_bias_data *
in
Pointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct.
dst
int16_t *
out
Pointer to the output matrix with "m" rows and "n" columns
dst_multipliers
const int32_t *
in
Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
dst_shifts
const int32_t *
in
Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)
lhs_rows
const int32_t
in
Number of LHS input rows
rhs_rows
const int32_t
in
Number of RHS input rows
rhs_cols
const int32_t
in
Number of LHS/RHS input columns
activation_min
const int32_t
in
Minimum value to clamp down the output. Range : int16
activation_max
const int32_t
in
Maximum value to clamp up the output. Range : int16
row_address_offset
const int32_t
in
Address offset between rows in output. NOTE: Only used for MVEI extension.
Returns
Returns of arm_nn_mat_mult_nt_t_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS` or `ARM_CMSIS_NN_NO_IMPL_ERROR` if not for MVE |---row_address_offset---| |____rhs_rows__________________|
| | |
| --- | --- |
| | |
| | |
| | | lhs_rows
| | |
| --- | --- |
| _______________ | ______________ |
Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in padded cases where the padding is -lhs_offset(Range: int8). Dimensions are the same for lhs and rhs.
Parameters
Parameters of arm_nn_depthwise_conv_nt_t_padded_s8
Name
Type
Direction
Description
lhs
const int8_t *
in
Input left-hand side matrix
rhs
const int8_t *
in
Input right-hand side matrix (transposed)
lhs_offset
const int32_t
in
LHS matrix offset(input offset). Range: -127 to 128
active_ch
const int32_t
in
Subset of total_ch processed
total_ch
const int32_t
in
Number of channels in LHS/RHS
out_shift
const int32_t *
in
Per channel output shift. Length of vector is equal to number of channels
out_mult
const int32_t *
in
Per channel output multiplier. Length of vector is equal to number of channels
out_offset
const int32_t
in
Offset to be added to the output values. Range: -127 to 128
activation_min
const int32_t
in
Minimum value to clamp the output to. Range: int8
activation_max
const int32_t
in
Maximum value to clamp the output to. Range: int8
row_x_col
const uint16_t
in
(row_dimension * col_dimension) of LHS/RHS matrix
output_bias
const int32_t *const
in
Per channel output bias. Length of vector is equal to number of channels
out
int8_t *
out
Output pointer
Returns
Returns of arm_nn_depthwise_conv_nt_t_padded_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` if an implementation is available or `ARM_CMSIS_NN_NO_IMPL_ERROR` otherwise
Necessary conditions of the planar rule that are cheap to test inline: at most 32 channels and stride 1. A caller can skip arm_nn_depthwise_conv_s8_planar() for layers that fail them without changing which layers it takes.
Parameters
Parameters of arm_nn_depthwise_conv_s8_planar_candidate
Name
Type
Direction
Description
dw_conv_params
const cmsis_nn_dw_conv_params *
in
Depthwise convolution parameters
input_dims
const cmsis_nn_dims *
in
Input tensor dimensions. Format: [1, H, W, C_IN]
Returns
Returns of arm_nn_depthwise_conv_s8_planar_candidate
Description
1 when the layer may take the planar path, 0 when it cannot.
The gate of armconvolves8smallcin(): upscaledims NULL, input depth 1 to 3 with filter depth equal to it, dilation 1, a kernel of at least 1x1 with kernel width…
The gate of arm_convolve_s8_small_cin(): upscale_dims NULL, input depth 1 to 3 with filter depth equal to it, dilation 1, a kernel of at least 1x1 with kernel width x depth at most 16 and at most 48 values, and a positive multiple of 4 output channels. Plain C; it evaluates the same on every build.
The gate of arm_convolve_s8_3x3_c16_s1(): upscale_dims NULL, input and filter depth 16, a 3x3 kernel, and stride and dilation 1. Plain C; it evaluates the same on every build.
The group check of arm_convolve_s8(), for its direct entries: with groups = C_IN / filter C, C_IN or C_OUT is not a multiple of groups. A filter C of zero or above C_IN gives no group count and is not reported.
Plane size in bytes that arm_nn_depthwise_conv_s8_planar() needs for a layer, or -1 when the layer is not one it takes. The rule is plain C and evaluates the same on every build.
Parameters
Parameters of arm_nn_depthwise_conv_s8_planar_bytes
s8 depthwise convolution with channel multiplier 1 and stride 1, vectorized across the output pixels of one channel plane instead of across channels. It serves the few-channel and 1xk layers of arm_depthwise_conv_s8_opt(), with the same scratch buffer and weight sums.
Parameters
Parameters of arm_nn_depthwise_conv_s8_planar
Name
Type
Direction
Description
ctx
const cmsis_nn_context *
in, out
Scratch buffer of `arm_depthwise_conv_s8_opt_get_buffer_size()` bytes
weight_sum_ctx
const cmsis_nn_context *
in
Per-channel weight sums from `arm_depthwise_convolve_weight_sum()`, bias included
`ARM_CMSIS_NN_SUCCESS` when the layer was computed, or `ARM_CMSIS_NN_NO_IMPL_ERROR` when it is not one this path takes or its plane does not fit in ctx->size (then nothing is written), or MVE is not available.
Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. rhs consists of packed int4 data. Dimensions are the same for lhs and rhs.
Parameters
Parameters of arm_nn_depthwise_conv_nt_t_s4
Name
Type
Direction
Description
lhs
const int8_t *
in
Input left-hand side matrix
rhs
const int8_t *
in
Input right-hand side matrix (transposed). Consists of int4 data packed in an int8 buffer.
lhs_offset
const int32_t
in
LHS matrix offset(input offset). Range: -127 to 128
active_ch
const int32_t
in
Subset of total_ch processed
total_ch
const int32_t
in
Number of channels in LHS/RHS
out_shift
const int32_t *
in
Per channel output shift. Length of vector is equal to number of channels.
out_mult
const int32_t *
in
Per channel output multiplier. Length of vector is equal to number of channels.
out_offset
const int32_t
in
Offset to be added to the output values. Range: -127 to 128
activation_min
const int32_t
in
Minimum value to clamp the output to. Range: int8
activation_max
const int32_t
in
Maximum value to clamp the output to. Range: int8
row_x_col
const uint16_t
in
(row_dimension * col_dimension) of LHS/RHS matrix
output_bias
const int32_t *const
in
Per channel output bias. Length of vector is equal to number of channels.
out
int8_t *
out
Output pointer
Returns
Returns of arm_nn_depthwise_conv_nt_t_s4
Description
The function returns one of the two
- Updated output pointer if an implementation is available
- NULL if no implementation is available.
Matrix-multiplication function for convolution with per-channel requantization and 4 bit weights.
This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.
Parameters
Parameters of arm_nn_mat_mult_kernel_s4_s16
Name
Type
Direction
Description
input_a
const int8_t *
in
pointer to operand A, int8 packed with 2x int4.
input_b
const int16_t *
in
pointer to operand B, always consists of 2 vectors.
output_ch
const uint16_t
in
number of rows of A
out_shift
const int32_t *
in
pointer to per output channel requantization shift parameter.
out_mult
const int32_t *
in
pointer to per output channel requantization multiplier parameter.
out_offset
const int32_t
in
output tensor offset.
activation_min
const int32_t
in
minimum value to clamp the output to. Range : int8
activation_max
const int32_t
in
maximum value to clamp the output to. Range : int8
num_col_a
const int32_t
in
number of columns of A
output_bias
const int32_t *const
in
per output channel bias. Range : int32
out_0
int8_t *
in, out
pointer to output
Returns
Returns of arm_nn_mat_mult_kernel_s4_s16
Description
The function returns one of the two
1. The incremented output pointer for a successful operation or
2. NULL if implementation is not available.
Matrix-multiplication function for convolution with per-channel requantization.
This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.
Parameters
Parameters of arm_nn_mat_mult_kernel_s8_s16
Name
Type
Direction
Description
input_a
const int8_t *
in
pointer to operand A
input_b
const int16_t *
in
pointer to operand B, always consists of 2 vectors.
output_ch
const uint16_t
in
number of rows of A
out_shift
const int32_t *
in
pointer to per output channel requantization shift parameter.
out_mult
const int32_t *
in
pointer to per output channel requantization multiplier parameter.
out_offset
const int32_t
in
output tensor offset.
activation_min
const int16_t
in
minimum value to clamp the output to. Range : int8
activation_max
const int16_t
in
maximum value to clamp the output to. Range : int8
num_col_a
const int32_t
in
number of columns of A
aligned_num_col_a
const int32_t
in
number of columns of A aligned by 4
output_bias
const int32_t *const
in
per output channel bias. Range : int32
out_0
int8_t *
in, out
pointer to output
Returns
Returns of arm_nn_mat_mult_kernel_s8_s16
Description
The function returns one of the two
1. The incremented output pointer for a successful operation or
2. NULL if implementation is not available.
Matrix-multiplication function for convolution with per-channel requantization, supporting an address offset between rows.
This function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max.
This function is slighly less performant than arm_nn_mat_mult_kernel_s8_s16, but allows support for grouped convolution. Supported framework: TensorFlow Lite micro.
Parameters
Parameters of arm_nn_mat_mult_kernel_row_offset_s8_s16
Name
Type
Direction
Description
input_a
const int8_t *
in
pointer to operand A
input_b
const int16_t *
in
pointer to operand B, always consists of 2 vectors.
output_ch
const uint16_t
in
number of rows of A
out_shift
const int32_t *
in
pointer to per output channel requantization shift parameter.
out_mult
const int32_t *
in
pointer to per output channel requantization multiplier parameter.
out_offset
const int32_t
in
output tensor offset.
activation_min
const int16_t
in
minimum value to clamp the output to. Range : int8
activation_max
const int16_t
in
maximum value to clamp the output to. Range : int8
num_col_a
const int32_t
in
number of columns of A
aligned_num_col_a
const int32_t
in
number of columns of A aligned by 4
output_bias
const int32_t *const
in
per output channel bias. Range : int32
row_address_offset
const int32_t
in
address offset between rows in the output
out_0
int8_t *
in, out
pointer to output
Returns
Returns of arm_nn_mat_mult_kernel_row_offset_s8_s16
Description
The function returns one of the two
1. The incremented output pointer for a successful operation or
2. NULL if implementation is not available.
Essentially returns (val * multiplier)/(2 ^ shift) with different rounding depending if CMSIS_NN_USE_SINGLE_ROUNDING is defined or not.
Parameters
Parameters of arm_nn_requantize
Name
Type
Direction
Description
val
const int32_t
in
Value to be requantized
multiplier
const int32_t
in
Multiplier. Range {NN_Q31_MIN + 1, Q32_MAX}
shift
const int32_t
in
Shift. Range: {-31, 30} Default branch: If shift is positive left shift 'val * multiplier' with shift If shift is negative right shift 'val * multiplier' with abs(shift) Single round branch: Input for total_shift in divide by '2 ^ total_shift'
Returns
Returns of arm_nn_requantize
Description
Default branch: Returns (val * multiplier) with rounding divided by (2 ^ shift) with rounding Single round branch: Returns (val * multiplier)/(2 ^ (31 - shift)) with rounding
Update LSTM function for an iteration step using s8 input and output, and s16 internally.
Parameters
Parameters of arm_nn_lstm_step_s8
Name
Type
Direction
Description
data_in
const int8_t *
in
Data input pointer
hidden_in
const int8_t *
in
Hidden state/ recurrent input pointer
hidden_out
int8_t *
out
Hidden state/ recurrent output pointer
params
const cmsis_nn_lstm_params *
in
Struct containg all information about the lstm operator, see arm_nn_types.
buffers
cmsis_nn_lstm_context *
in, out
Struct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types.
batch_offset
const int32_t
in
Number of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps.
Update LSTM function for an iteration step using s16 input and output, and s16 internally.
Parameters
Parameters of arm_nn_lstm_step_s16
Name
Type
Direction
Description
data_in
const int16_t *
in
Data input pointer
hidden_in
const int16_t *
in
Hidden state/ recurrent input pointer
hidden_out
int16_t *
out
Hidden state/ recurrent output pointer
params
const cmsis_nn_lstm_params *
in
Struct containg all information about the lstm operator, see arm_nn_types.
buffers
cmsis_nn_lstm_context *
in, out
Struct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types.
batch_offset
const int32_t
in
Number of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps.
The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a “batched” vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).
Parameters
Parameters of arm_nn_vec_mat_mul_result_acc_s8_s16
Name
Type
Direction
Description
lhs
const int8_t *
in
Batched vector
rhs
const int8_t *
in
Weights - input matrix (H(Rows)xW(Columns))
effective_bias
const int32_t *
in
Bias + lhs_offset * kernel_sum term precalculated into a constant vector.
dst
int16_t *
out
Output
dst_multiplier
const int32_t
in
Multiplier for quantization
dst_shift
const int32_t
in
Shift for quantization
rhs_cols
const int32_t
in
Vector/matarix column length
rhs_rows
const int32_t
in
Row count of matrix
batches
const int32_t
in
Batch size
batch_offset
const int32_t
in
Number of timesteps between consecutive batches in input, see arm_nn_lstm_step_s8. Note that the output is always stored with sequential batches.
The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a “batched” vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).
Parameters
Parameters of arm_nn_vec_mat_mul_result_acc_s16
Name
Type
Direction
Description
lhs
const int16_t *
in
Batched vector
rhs
const int8_t *
in
Weights - input matrix (H(Rows)xW(Columns))
effective_bias
const int64_t *
in
Bias + lhs_offset * kernel_sum term precalculated into a constant vector.
dst
int16_t *
out
Output
dst_multiplier
const int32_t
in
Multiplier for quantization
dst_shift
const int32_t
in
Shift for quantization
rhs_cols
const int32_t
in
Vector/matarix column length
rhs_rows
const int32_t
in
Row count of matrix
batches
const int32_t
in
Batch size
batch_offset
const int32_t
in
Number of timesteps between consecutive batches in input, see arm_nn_lstm_step_s16. Note that the output is always stored with sequential batches.
Number of timesteps between consecutive batches in output, see arm_nn_lstm_step_s8. Note that it is assumed that the input is stored with sequential batches.
Parameters of arm_elementwise_mul_s16_batch_offset
Name
Type
Direction
Description
input_1_vect
const int16_t *
in
pointer to input vector 1
input_2_vect
const int16_t *
in
pointer to input vector 2
output
int16_t *
in, out
pointer to output vector
out_offset
const int32_t
in
output offset
out_mult
const int32_t
in
output multiplier
out_shift
const int32_t
in
output shift
block_size
const int32_t
in
number of samples per batch
batch_size
const int32_t
in
number of samples per batch
batch_offset
const int32_t
in
Number of timesteps between consecutive batches in output, see arm_nn_lstm_step_s16. Note that it is assumed that the input is stored with sequential batches.
Reports whether the reduced axes of a 4-D tensor form one contiguous block followed by kept axes, as in a NHWC mean over H and W, and gives the flattened sizes.
Reports whether the reduced axes of a 4-D tensor form one contiguous block followed by kept axes, as in a NHWC mean over H and W, and gives the flattened sizes. Axes of size 1 are ignored.
Parameters
Parameters of arm_reduce_get_middle_block_from_arrays
Name
Type
Direction
Description
in_dims
const int32_t
in
4-element array {n, h, w, c}
axis_arr
const int32_t
in
4-element mask {axis_n, axis_h, axis_w, axis_c}
outer
int32_t *
out
Product of the dims before the reduced block
reduce
int32_t *
out
Product of the reduced dims
inner
int32_t *
out
Product of the dims after the reduced block
Returns
Returns of arm_reduce_get_middle_block_from_arrays
Description
1 if the input is [outer, reduce, inner] with the middle dim reduced and inner > 1, otherwise 0
One element of armsqrts16tablefree(): the float32 chain the MVE path evaluates per lane, so the two agree bit for bit on any IEEE-754 float32 implementation wi…
One element of arm_sqrt_s16_tablefree(): the float32 chain the MVE path evaluates per lane, so the two agree bit for bit on any IEEE-754 float32 implementation with round-to-nearest-even and a fused multiply-add (fmaf). Every product after the pre-scale either has two uses or feeds an fmaf or a conversion, never another lone multiply, so a compiler allowed to reassociate (-ffast-math) still has no chain to reorder, and no product feeds a bare add, so there is nothing to contract.
Parameters
Parameters of arm_nn_sqrt_s16_tablefree_element
Name
Type
Direction
Description
value
const int32_t
in
input code; values <= 0 give 0
scale
const float
in
input_scale / (output_scale * output_scale) as float32