Skip to content
heliaCORE
API reference
HELIA HUB

arm_nnfunctions.h

Machine-readable model

function

s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.

Include/arm_nnfunctions.h:97

arm_cmsis_nn_status arm_convolve_wrapper_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.

Parameters of arm_convolve_wrapper_s4
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_wrapper_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
filter_dataconst int8_t *inFilter data pointer. Data type: int8 packed with 2x int4
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_wrapper_s4
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Get the required buffer size for armconvolvewrappers4.

Include/arm_nnfunctions.h:133

int32_t arm_convolve_wrapper_s4_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required buffer size for arm_convolve_wrapper_s4.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_convolve_wrapper_s4_get_buffer_size
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_wrapper_s4_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid.
function

Get the required buffer size for armconvolvewrappers4 for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:149

int32_t arm_convolve_wrapper_s4_get_buffer_size_mve(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required buffer size for arm_convolve_wrapper_s4 for Arm(R) Helium Architecture case.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_convolve_wrapper_s4_get_buffer_size_mve
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_wrapper_s4_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid.
function

Get the required buffer size for armconvolvewrappers4 for processors with DSP extension.

Include/arm_nnfunctions.h:164

int32_t arm_convolve_wrapper_s4_get_buffer_size_dsp(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required buffer size for arm_convolve_wrapper_s4 for processors with DSP extension.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_convolve_wrapper_s4_get_buffer_size_dsp
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_wrapper_s4_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid.
function

s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.

Include/arm_nnfunctions.h:225

arm_cmsis_nn_status arm_convolve_wrapper_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.

  • On builds with ARM_MATH_MVEI (without ARM_MATH_AUTOVECTORIZE), a layer that would run arm_convolve_s8() and is in the gate of arm_convolve_s8_small_cin() or arm_convolve_s8_3x3_c16_s1() runs that entry instead, with the same result, scratch and weight sums. The input depth is checked first, so other layers skip both gates.
Parameters of arm_convolve_wrapper_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_wrapper_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
weight_sum_ctxconst cmsis_nn_context *inPer-output-channel weight sums, supplied by the caller. The selected kernel only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build: this wrapper dispatches to `arm_convolve_s8()`, `arm_convolve_1x1_s8()`, `arm_convolve_1x1_s8_fast()`, `arm_convolve_1_x_n_s8()` and `arm_convolve_1x1_out_s8()`. The buffer contents are consumed only on builds with the MVE extension (ARM_MATH_MVEI), and on those builds every one of those kernels diagnoses a NULL buf with ARM_CMSIS_NN_ARG_ERROR; on other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed that way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_wrapper_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Get the required buffer size for armconvolvewrappers8.

Include/arm_nnfunctions.h:264

int32_t arm_convolve_wrapper_s8_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required buffer size for arm_convolve_wrapper_s8.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_convolve_wrapper_s8_get_buffer_size
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_wrapper_s8_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid.
function

Get the required buffer size for armconvolves8 for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:278

int32_t arm_convolve_s8_get_buffer_size_mve(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for arm_convolve_s8 for Arm(R) Helium Architecture case.

The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas, so a shape whose dimension product overflows is only reported by the leg that actually computes a buffer for it.

Parameters of arm_convolve_s8_get_buffer_size_mve
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
Returns of arm_convolve_s8_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Get the required buffer size for armconvolvewrappers8 for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:290

int32_t arm_convolve_wrapper_s8_get_buffer_size_mve(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required buffer size for arm_convolve_wrapper_s8 for Arm(R) Helium Architecture case.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_convolve_wrapper_s8_get_buffer_size_mve
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_wrapper_s8_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid.
function

Get the required buffer size for armconvolvewrappers8 for processors with DSP extension.

Include/arm_nnfunctions.h:305

int32_t arm_convolve_wrapper_s8_get_buffer_size_dsp(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required buffer size for arm_convolve_wrapper_s8 for processors with DSP extension.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_convolve_wrapper_s8_get_buffer_size_dsp
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_wrapper_s8_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid.
function

s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.

Include/arm_nnfunctions.h:338

arm_cmsis_nn_status arm_convolve_wrapper_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.

Parameters of arm_convolve_wrapper_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_wrapper_s8_get_buffer_size will return the buffer_size if required The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst cmsis_nn_bias_data *inStruct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_convolve_wrapper_s16
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

s16 grouped convolution optimized for the case where filterdims->c == 1 and inputch == outputch (channel multiplier = 1).

Include/arm_nnfunctions.h:368

arm_cmsis_nn_status arm_convolve_s16_group_ch_mult_1(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

s16 grouped convolution optimized for the case where filter_dims->c == 1 and input_ch == output_ch (channel multiplier = 1).

Parameters of arm_convolve_s16_group_ch_mult_1
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context (unused, pass NULL-initialised).
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (strides, dilations, pads, activation).
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info (multiplier and shift).
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int16_t *inInput data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, 1]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions (unused, may be zero-initialised).
bias_dataconst cmsis_nn_bias_data *inOptional bias struct (int32 or int64). May be NULL.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_convolve_s16_group_ch_mult_1
Description
ARM_CMSIS_NN_SUCCESS on success.
function

Get the required buffer size for armconvolvewrappers16.

Include/arm_nnfunctions.h:398

int32_t arm_convolve_wrapper_s16_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required buffer size for arm_convolve_wrapper_s16.

An out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size.

Parameters of arm_convolve_wrapper_s16_get_buffer_size
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_wrapper_s16_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Get the required buffer size for armconvolvewrappers16 for for processors with DSP extension.

Include/arm_nnfunctions.h:412

int32_t arm_convolve_wrapper_s16_get_buffer_size_dsp(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required buffer size for arm_convolve_wrapper_s16 for for processors with DSP extension.

An out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size.

Parameters of arm_convolve_wrapper_s16_get_buffer_size_dsp
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_wrapper_s16_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Get the required buffer size for armconvolvewrappers16 for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:426

int32_t arm_convolve_wrapper_s16_get_buffer_size_mve(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required buffer size for arm_convolve_wrapper_s16 for Arm(R) Helium Architecture case.

An out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size.

Parameters of arm_convolve_wrapper_s16_get_buffer_size_mve
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_wrapper_s16_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Basic s4 convolution function.

Include/arm_nnfunctions.h:458

arm_cmsis_nn_status arm_convolve_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Basic s4 convolution function.

  1. Supported framework: TensorFlow Lite micro
  2. Additional memory is required for optimization. Refer to argument ‘ctx’ for details.
Parameters of arm_convolve_s4
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
filter_dataconst int8_t *inPacked Filter data pointer. Data type: int8 packed with 2x int4
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_s4
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Basic s4 convolution function with a requirement of even number of kernels.

Include/arm_nnfunctions.h:499

arm_cmsis_nn_status arm_convolve_even_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Basic s4 convolution function with a requirement of even number of kernels.

  1. Supported framework: TensorFlow Lite micro
  2. Additional memory is required for optimization. Refer to argument ‘ctx’ for details.
Parameters of arm_convolve_even_s4
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_even_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions. Note the product must be even.
filter_dataconst int8_t *inPacked Filter data pointer. Data type: int8 packed with 2x int4
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_even_s4
Description
The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` if not for MVE
function

Basic s8 convolution function.

Include/arm_nnfunctions.h:563

arm_cmsis_nn_status arm_convolve_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *upscale_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Basic s8 convolution function.

  1. Supported framework: TensorFlow Lite micro
  2. Additional memory is required for optimization. Refer to argument ‘ctx’ for details.
Parameters of arm_convolve_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
weight_sum_ctxconst cmsis_nn_context *inPer-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. For grouped convolution the entries run over all output_dims->c channels, groups laid out consecutively. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI); an unfilled buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, and a NULL buf is currently reported as ARM_CMSIS_NN_ARG_ERROR. On other builds this function currently derives the same quantity itself and does not read the context. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, CK] where HK, WK and CK are the spatial filter dimensions. CK != C_IN is used for grouped convolution, in which case the required conditions are C_IN = N * CK and C_OUT = N * M for N groups of size M.
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
upscale_dimsconst cmsis_nn_dims *inUpscale tensor dimensions for transpose. Format: [H_UP, W_UP]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR`
function

s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model.

Include/arm_nnfunctions.h:620

arm_cmsis_nn_status arm_convolve_s8_small_cin(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *upscale_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model. It copies each kernel row with one predicated vector load and multiplies four output channels per step.

  • The output is identical to arm_convolve_s8(). The bias is read through the weight sums, which arm_convolve_weight_sum() fills as for arm_convolve_s8(); bias_dims and bias_data are unused.
  • Gate: upscale_dims NULL, C_IN from 1 to 3 with CK equal to C_IN (one group), dilation 1 in both dimensions, WK and HK at least 1 with WK x C_IN at most 16 and HK x WK x C_IN at most 48, and C_OUT a positive multiple of 4. Stride, padding and batch count are as for arm_convolve_s8().
  • Scratch: ctx->buf holds arm_convolve_s8_get_buffer_size() bytes (4 x 16 x ceil(HK x WK x C_IN / 16) on ARM_MATH_MVEI builds), the same as arm_convolve_s8(), and needs no alignment.
  • It is a direct entry: arm_convolve_s8() does not call it, and arm_convolve_wrapper_s8() calls it for layers in the gate that it would otherwise pass to arm_convolve_s8(). A caller that selects the kernel per layer ahead of time calls it for layers in the gate and arm_convolve_s8() for every other layer, or on ARM_CMSIS_NN_NO_IMPL_ERROR. Both take the same arguments, scratch and weight sums.
Parameters of arm_convolve_s8_small_cin
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context with `arm_convolve_s8_get_buffer_size()` bytes of scratch, all of which may be written
weight_sum_ctxconst cmsis_nn_context *inPer-output-channel weight sums, as for `arm_convolve_s8()`
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters, as for `arm_convolve_s8()`
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, CK]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
upscale_dimsconst cmsis_nn_dims *inUpscale tensor dimensions for transpose. Format: [H_UP, W_UP]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_s8_small_cin
Description
The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - an argument error that `arm_convolve_s8()` reports: ctx->buf is NULL, C_IN or C_OUT is not a multiple of the group count C_IN / CK, or weight_sum_ctx->buf is NULL on builds with ARM_MATH_MVEI. These are checked before the gate. `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer is outside the gate below, or the build lacks ARM_MATH_MVEI or defines ARM_MATH_AUTOVECTORIZE; nothing is written, to the output or to the scratch `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

s8 3x3 convolution over 16 input channels with unit stride.

Include/arm_nnfunctions.h:672

arm_cmsis_nn_status arm_convolve_s8_3x3_c16_s1(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *upscale_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 3x3 convolution over 16 input channels with unit stride. It reads the kernel rows of a patch inside the input in place, copying only patches that cross the border, and multiplies four output pixels per filter load.

  • The output is identical to arm_convolve_s8(). The bias is read through the weight sums; bias_dims and bias_data are unused.
  • Gate: upscale_dims NULL, C_IN and CK both 16 (one group), HK and WK both 3, and stride and dilation 1 in both dimensions. Padding, batch count and C_OUT are as for arm_convolve_s8().
  • It is a direct entry: arm_convolve_s8() does not call it, and arm_convolve_wrapper_s8() calls it for layers in the gate that it would otherwise pass to arm_convolve_s8(). A caller that selects the kernel per layer ahead of time calls it for layers in the gate and arm_convolve_s8() for every other layer, or on ARM_CMSIS_NN_NO_IMPL_ERROR. The gate does not overlap that of arm_convolve_s8_small_cin().
Parameters of arm_convolve_s8_3x3_c16_s1
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context with `arm_convolve_s8_get_buffer_size()` bytes of scratch (576 bytes on ARM_MATH_MVEI builds), all of which may be written
weight_sum_ctxconst cmsis_nn_context *inPer-output-channel weight sums, as for `arm_convolve_s8()`
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters, as for `arm_convolve_s8()`
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, CK]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
upscale_dimsconst cmsis_nn_dims *inUpscale tensor dimensions for transpose. Format: [H_UP, W_UP]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_s8_3x3_c16_s1
Description
The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_convolve_s8_small_cin()` `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer is outside the gate below, or the build lacks ARM_MATH_MVEI or defines ARM_MATH_AUTOVECTORIZE; nothing is written, to the output or to the scratch `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

Get the required buffer size for s4 convolution function.

Include/arm_nnfunctions.h:698

int32_t arm_convolve_s4_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for s4 convolution function.

The dimensions and the byte count are both checked here, so an out-of-range shape returns -1 on every build target rather than a wrapped size.

Parameters of arm_convolve_s4_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
Returns of arm_convolve_s4_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Get the required buffer size for armconvolveevens4.

Include/arm_nnfunctions.h:713

int32_t arm_convolve_even_s4_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for arm_convolve_even_s4.

Forwards to arm_convolve_s4_get_buffer_size(): the even_s4 kernel stages up to four im2col rows of filter_dims->w * filter_dims->h * input_dims->c int8 elements, byte-for-byte the size that sizer returns. The equality, including the -1 answers for out-of-range shapes, is pinned by a Unity test.

Parameters of arm_convolve_even_s4_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
Returns of arm_convolve_even_s4_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Get the required buffer size for s8 convolution function.

Include/arm_nnfunctions.h:729

int32_t arm_convolve_s8_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for s8 convolution function.

The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas, so a shape whose dimension product overflows is only reported by the leg that actually computes a buffer for it.

Parameters of arm_convolve_s8_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
Returns of arm_convolve_s8_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Get the required buffer size for s8 convolution and depthwise convolution weight sum.

Include/arm_nnfunctions.h:742

int32_t arm_convolve_s8_get_weights_sum_size(const cmsis_nn_dims *output_dims)

Get the required buffer size for s8 convolution and depthwise convolution weight sum.

For a valid (non-negative, in-range) output_dims->c, returns output_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. A negative or out-of-range output_dims->c returns -1 on builds with the MVE extension; elsewhere no weight sum buffer is used and the answer stays 0.

Parameters of arm_convolve_s8_get_weights_sum_size
NameTypeDirectionDescription
output_dimsconst cmsis_nn_dims *inOutput (activation) tensor dimensions. Format: [N, H, W, C_COUT]
Returns of arm_convolve_s8_get_weights_sum_size
Description
The function returns required weight sum buffer size in bytes, or -1 if output_dims->c is negative or the required size would not fit in an int32_t
function

Wrapper to select optimal transposed convolution algorithm depending on parameters.

Include/arm_nnfunctions.h:812

arm_cmsis_nn_status arm_transpose_conv_wrapper_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_context *reverse_conv_ctx,
const cmsis_nn_transpose_conv_params *transpose_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Wrapper to select optimal transposed convolution algorithm depending on parameters.

  1. Supported framework: TensorFlow Lite micro
  2. Additional memory is required for optimization. Refer to arguments ‘ctx’ and ‘reverse_conv_ctx’ for details.
  3. Dilation is not supported: transpose_conv_params->dilation must be 1 in both dimensions, otherwise ARM_CMSIS_NN_ARG_ERROR is returned.
Parameters of arm_transpose_conv_wrapper_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. `arm_transpose_conv_s8_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
weight_sum_ctxconst cmsis_nn_context *inPer-output-channel weight sums, supplied by the caller. The function only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and transpose_conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing transpose_conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Compute the sums over filter_data exactly as passed to this function: this wrapper guarantees that whatever filter preparation it performs internally preserves the per-output-channel sums, so no reversed or otherwise rearranged copy of the weights is needed for this step. Pass a valid context on every build. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the reverse-convolution route forwards this context to `arm_convolve_s8()`; an unfilled buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, and a NULL buf is currently reported as ARM_CMSIS_NN_ARG_ERROR. On other builds the contents are currently not read. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.
reverse_conv_ctxconst cmsis_nn_context *in, outFunction context for the reversed filter used when this wrapper routes to the reverse convolution. Holds filter height * filter width * input channels * output channels int8 values; `arm_transpose_conv_s8_get_reverse_conv_buffer_size()` returns the required size (0 when the reverse-convolution route is not taken). The caller is expected to clear the buffer, if applicable, for security reasons.
transpose_conv_paramsconst cmsis_nn_transpose_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of transpose_conv_params->input_offset : [-127, 128] Range of transpose_conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each out channel.
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_transpose_conv_wrapper_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Basic s8 transpose convolution function.

Include/arm_nnfunctions.h:864

arm_cmsis_nn_status arm_transpose_conv_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_transpose_conv_params *transpose_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Basic s8 transpose convolution function.

  1. Supported framework: TensorFlow Lite micro
  2. Additional memory is required for optimization. Refer to argument ‘ctx’ for details; ‘output_ctx’ is unused.
  3. Dilation is not supported: transpose_conv_params->dilation must be 1 in both dimensions, otherwise ARM_CMSIS_NN_ARG_ERROR is returned.
Parameters of arm_transpose_conv_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_transpose_conv_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
output_ctxconst cmsis_nn_context *in, outNot accessed by this function: its buffer is neither read nor written, and it therefore has no size requirement. The parameter exists only to keep one signature across the transpose-conv family, whose float twins ignore it the same way; `arm_transpose_conv_wrapper_s8()` forwards its reverse_conv_ctx into this slot. In-tree callers pass a valid context, whose buf may be NULL.
transpose_conv_paramsconst cmsis_nn_transpose_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of transpose_conv_params->input_offset : [-127, 128] Range of transpose_conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each out channel.
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_transpose_conv_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Get the required buffer size for ctx in s8 transpose conv function.

Include/arm_nnfunctions.h:894

int32_t arm_transpose_conv_s8_get_buffer_size(
const cmsis_nn_transpose_conv_params *transposed_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *out_dims
)

Get the required buffer size for ctx in s8 transpose conv function.

The returned size is safe for both arm_transpose_conv_s8() and arm_transpose_conv_wrapper_s8(): it is the larger of the two routes’ requirements, so it may exceed what the wrapper’s reverse-convolution route alone would need. When either route is out of range the sentinel is propagated ahead of that comparison, so -1 is never collapsed into a plausible positive size by the other route.

Parameters of arm_transpose_conv_s8_get_buffer_size
NameTypeDirectionDescription
transposed_conv_paramsconst cmsis_nn_transpose_conv_params *inTransposed convolution parameters
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
out_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_transpose_conv_s8_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative, either stride is not positive, or the required size would not fit in an int32_t
function

Get the required buffer size for outputctx in s8 transpose conv function.

Include/arm_nnfunctions.h:910

int32_t arm_transpose_conv_s8_get_reverse_conv_buffer_size(
const cmsis_nn_transpose_conv_params *transposed_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims
)

Get the required buffer size for output_ctx in s8 transpose conv function.

Parameters of arm_transpose_conv_s8_get_reverse_conv_buffer_size
NameTypeDirectionDescription
transposed_conv_paramsconst cmsis_nn_transpose_conv_params *inTransposed convolution parameters
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
Returns of arm_transpose_conv_s8_get_reverse_conv_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Get size of additional buffer required by armtransposeconvs8() for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:923

int32_t arm_transpose_conv_s8_get_buffer_size_mve(
const cmsis_nn_transpose_conv_params *transposed_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *out_dims
)

Get size of additional buffer required by arm_transpose_conv_s8() for Arm(R) Helium Architecture case.

The returned size is safe for both arm_transpose_conv_s8() and arm_transpose_conv_wrapper_s8(): it is the larger of the two routes’ requirements, so it may exceed what the wrapper’s reverse-convolution route alone would need. When either route is out of range the sentinel is propagated ahead of that comparison, so -1 is never collapsed into a plausible positive size by the other route.

Parameters of arm_transpose_conv_s8_get_buffer_size_mve
NameTypeDirectionDescription
transposed_conv_paramsconst cmsis_nn_transpose_conv_params *inTransposed convolution parameters
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
out_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_transpose_conv_s8_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative, either stride is not positive, or the required size would not fit in an int32_t
function

Basic s16 convolution function.

Include/arm_nnfunctions.h:958

arm_cmsis_nn_status arm_convolve_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

Basic s16 convolution function.

  1. Supported framework: TensorFlow Lite micro
  2. Additional memory is required for optimization. Refer to argument ‘ctx’ for details.
Parameters of arm_convolve_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst cmsis_nn_bias_data *inStruct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_convolve_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR`
function

Pointwise s16 convolution function: no stride, no padding, no dilation.

Include/arm_nnfunctions.h:999

arm_cmsis_nn_status arm_convolve_1x1_s16_ns_np_nd(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

Pointwise s16 convolution function: no stride, no padding, no dilation.

  1. Supported framework: TensorFlow Lite micro
Parameters of arm_convolve_1x1_s16_ns_np_nd
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst cmsis_nn_bias_data *inStruct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_convolve_1x1_s16_ns_np_nd
Description
The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR`
function

armconvolves16fastsmallkernel function.

Include/arm_nnfunctions.h:1040

arm_cmsis_nn_status arm_convolve_s16_fast_small_kernel(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const cmsis_nn_bias_data *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

arm_convolve_s16_fast_small_kernel function. The kernel size is <=8

  1. Supported framework: TensorFlow Lite micro
Parameters of arm_convolve_s16_fast_small_kernel
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst cmsis_nn_bias_data *inStruct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_convolve_s16_fast_small_kernel
Description
The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR`
function

Get the required buffer size for s16 convolution function.

Include/arm_nnfunctions.h:1065

int32_t arm_convolve_s16_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for s16 convolution function.

The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas.

Parameters of arm_convolve_s16_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions
Returns of arm_convolve_s16_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Fast s4 version for 1x1 convolution (non-square shape).

Include/arm_nnfunctions.h:1098

arm_cmsis_nn_status arm_convolve_1x1_s4_fast(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Fast s4 version for 1x1 convolution (non-square shape).

  • Supported framework : TensorFlow Lite Micro

  • The following constrains on the arguments apply

    1. conv_params->padding.w = conv_params->padding.h = 0
    2. conv_params->stride.w = conv_params->stride.h = 1
Parameters of arm_convolve_1x1_s4_fast
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_1x1_s4_fast_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, 1, 1, C_IN]
filter_dataconst int8_t *inFilter data pointer. Data type: int8 packed with 2x int4
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_1x1_s4_fast
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

s4 version for 1x1 convolution with support for non-unity stride values

Include/arm_nnfunctions.h:1138

arm_cmsis_nn_status arm_convolve_1x1_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s4 version for 1x1 convolution with support for non-unity stride values

  • Supported framework : TensorFlow Lite Micro

  • The following constrains on the arguments apply

    1. conv_params->padding.w = conv_params->padding.h = 0
Parameters of arm_convolve_1x1_s4
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. None is required by this function.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, 1, 1, C_IN]
filter_dataconst int8_t *inFilter data pointer. Data type: int8 packed with 2x int4
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_1x1_s4
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Fast s8 version for 1x1 convolution (non-square shape).

Include/arm_nnfunctions.h:1204

arm_cmsis_nn_status arm_convolve_1x1_s8_fast(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Fast s8 version for 1x1 convolution (non-square shape).

  • Supported framework : TensorFlow Lite Micro

  • The following constrains on the arguments apply

    1. conv_params->padding.w = conv_params->padding.h = 0
    2. conv_params->stride.w = conv_params->stride.h = 1
Parameters of arm_convolve_1x1_s8_fast
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_1x1_s8_fast_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
weight_sum_ctxconst cmsis_nn_context *inPer-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. This function reads the buffer contents only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. On other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. Note also that on an Arm Compiler build (__ARMCC_VERSION >= 6010050) with ARM_MATH_DSP and without ARM_MATH_MVEI, supplying ctx->buf selects a buffered path that never reads weight_sum_ctx. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, 1, 1, C_IN]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_1x1_s8_fast
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Get the required buffer size for armconvolve1x1s4fast.

Include/arm_nnfunctions.h:1225

int32_t arm_convolve_1x1_s4_fast_get_buffer_size(const cmsis_nn_dims *input_dims)

Get the required buffer size for arm_convolve_1x1_s4_fast.

Parameters of arm_convolve_1x1_s4_fast_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions
Returns of arm_convolve_1x1_s4_fast_get_buffer_size
Description
The function returns the required buffer size in bytes, or -1 if input_dims->c is negative. No build needs this scratch buffer, so every valid shape returns 0.
function

Get the required buffer size for armconvolve1x1s8fast.

Include/arm_nnfunctions.h:1236

int32_t arm_convolve_1x1_s8_fast_get_buffer_size(const cmsis_nn_dims *input_dims)

Get the required buffer size for arm_convolve_1x1_s8_fast.

Parameters of arm_convolve_1x1_s8_fast_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) dimensions
Returns of arm_convolve_1x1_s8_fast_get_buffer_size
Description
The function returns the required buffer size in bytes, or -1 if input_dims->c is negative. On builds that need this scratch buffer it also returns -1 if the required size would not fit in an int32_t; other builds need no buffer and return 0.
function

s8 version for 1x1 convolution with support for non-unity stride values

Include/arm_nnfunctions.h:1285

arm_cmsis_nn_status arm_convolve_1x1_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 version for 1x1 convolution with support for non-unity stride values

  • Supported framework : TensorFlow Lite Micro

  • The following constrains on the arguments apply

    1. conv_params->padding.w = conv_params->padding.h = 0
Parameters of arm_convolve_1x1_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. None is required by this function.
weight_sum_ctxconst cmsis_nn_context *inPer-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. This function reads the buffer contents only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. On other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, 1, 1, C_IN]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_1x1_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

1xn convolution

Include/arm_nnfunctions.h:1357

arm_cmsis_nn_status arm_convolve_1_x_n_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

1xn convolution

  • Supported framework : TensorFlow Lite Micro

  • The following constraints on the arguments apply

    1. input_dims->h, filter_dims->h and output_dims->h equal 1, and conv_params->padding.h is 0
    2. conv_params->dilation.w is 1 and conv_params->stride.w is positive
    3. conv_params->stride.w * input_dims->c is a multiple of 4
    4. conv_params->padding.w, input_dims->w and output_dims->w are not negative, and filter_dims->w is at least 1
  • Any horizontal padding and output width are handled, including an odd total padding, a filter wider than the input and a VALID layer whose stride leaves trailing input unused. On MVE builds the output columns whose window starts before or ends past the input read a padded copy of the input columns they span, staged in ctx; the other columns read the input in place.

Parameters of arm_convolve_1_x_n_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_1_x_n_s8_get_buffer_size will return the buffer_size if required. buf must not be NULL. On builds with the MVE extension (ARM_MATH_MVEI) a non-zero ctx->size smaller than the staging the layer needs is rejected with ARM_CMSIS_NN_ARG_ERROR. The caller is expected to clear the buffer, if applicable, for security reasons.
weight_sum_ctxconst cmsis_nn_context *inPer-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. The contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR; on other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_1_x_n_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Pre-computes per-output-channel weight sums for a standard convolution.

Include/arm_nnfunctions.h:1409

arm_cmsis_nn_status arm_convolve_weight_sum(
int32_t *vector_sum_buf,
const int8_t *rhs,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
const int32_t lhs_offset,
const int32_t *bias_data
)

Pre-computes per-output-channel weight sums for a standard convolution.

  • Supported framework : TensorFlow Lite Micro
  • The buffer pointed to by vector_sum_buf must be at least output_dims->c × sizeof(int32_t) bytes. arm_convolve_s8_get_weights_sum_size() returns that size on builds that use the sums, 0 elsewhere, and -1 for an output_dims->c that is negative or too large to size.
  • Layout: one int32 per output channel, indexed 0..output_dims->c - 1. Entry j holds lhs_offset * sum(weights of output channel j) + bias_data[j], i.e. the bias and the input-offset contribution folded together. For grouped convolution the entries run over all output channels, with the groups laid out consecutively.
  • This is the buffer the weight_sum_ctx parameter of the s8 convolution kernels carries. Those kernels currently treat it as an input they only read, so it has to be filled before the call - see the individual functions for what each one currently does on MVE and non-MVE builds.
  • Reuse and invalidation: the contents depend only on rhs, bias_data and lhs_offset. They do not depend on the activations, so a buffer stays valid across calls and across batches for as long as those three are unchanged - for a static model the sums can be computed once at load time rather than per inference. Recompute whenever the weights, the bias or the input offset change (for example on requantization or a weight reload). The buffer is sized by one layer’s output_dims->c and is specific to that layer’s weights, so it cannot be shared between layers; give each layer its own.
  • Returns ARM_CMSIS_NN_NO_IMPL_ERROR on builds without the MVE extension, where the sums are currently not consumed.
Parameters of arm_convolve_weight_sum
NameTypeDirectionDescription
vector_sum_bufint32_t *outPointer to the buffer that will hold the weight sums.
rhsconst int8_t *inPointer to the filter weights. Data type: int8
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, KH, KW, C_IN]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
lhs_offsetconst int32_tinInput-offset added to every input element before MAC. Range: [-127, 128]
bias_dataconst int32_t *inOptional bias pointer. Data type: int32
Returns of arm_convolve_weight_sum
Description
`ARM_CMSIS_NN_ARG_ERROR` on invalid arguments, `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed and the buffer is left untouched, or `ARM_CMSIS_NN_SUCCESS` on success. Portable code should not treat the `NO_IMPL_ERROR` case as a failure.
function

Pre-computes per-channel weight sums for a depthwise convolution.

Include/arm_nnfunctions.h:1458

arm_cmsis_nn_status arm_depthwise_convolve_weight_sum(
int32_t *vector_sum_buf,
int8_t *scratch_buf,
const int8_t *rhs,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
const int32_t lhs_offset,
const int32_t *bias_data
)

Pre-computes per-channel weight sums for a depthwise convolution.

  • Supported framework : TensorFlow Lite Micro
  • Layout: one int32 per channel, sized by arm_convolve_s8_get_weights_sum_size(). Entry j holds bias_data[j] + lhs_offset * sum(kernel values of channel j).
  • Reuse and invalidation follow the same rules as arm_convolve_weight_sum(): the contents depend only on rhs, bias_data and lhs_offset, so they may be computed once and reused until one of those changes, and they are specific to a single layer.
  • Returns ARM_CMSIS_NN_NO_IMPL_ERROR on builds without the MVE extension, where the sums are currently not consumed.
  • Not interchangeable with arm_convolve_weight_sum(): this function walks the channel-interleaved depthwise layout [1, KH, KW, C_OUT] with a stride of C_OUT, whereas arm_convolve_weight_sum() sums contiguous runs of KH * KW * C_IN weights. The two agree only by coincidence. Several in-tree tests do fill a depthwise weight_sum_ctx with arm_convolve_weight_sum() and are still correct, for one of three unrelated reasons: arm_depthwise_conv_wrapper_s8() does not consume the buffer on that route at all (ch_mult != 1, batches != 1, or a dilation the optimized route does not take - see that function); the wrapper converts the layer to a regular convolution, so conv-style sums are what is wanted; or C_OUT is 1, which collapses the stride-C_OUT walk to a contiguous one and makes the two helpers compute identical values. None of those generalise, so do not read them as licence to substitute one helper for the other. Use this function wherever the sums are actually read.
Parameters of arm_depthwise_convolve_weight_sum
NameTypeDirectionDescription
vector_sum_bufint32_t *outBuffer to hold the computed weight sums.
scratch_bufint8_t *in, outCurrently unused: the implementation does not read or write it on any build, so NULL is accepted. Retained for signature compatibility; if a real buffer is passed, the caller is expected to clear it for security reasons.
rhsconst int8_t *inDepthwise convolution weights. Data type: int8
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise-convolution parameters (stride, dilation, pad, etc.)
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, KH, KW, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
lhs_offsetconst int32_tinInput-offset applied before MAC. Range: [-127, 128]
bias_dataconst int32_t *inOptional bias pointer. Data type: int32
Returns of arm_depthwise_convolve_weight_sum
Description
`ARM_CMSIS_NN_ARG_ERROR` on invalid arguments, `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed and the buffer is left untouched, or `ARM_CMSIS_NN_SUCCESS` on success. Portable code should not treat the `NO_IMPL_ERROR` case as a failure.
function

Optimised convolution for 1x1 output images (shape of BX1x1xCOUT) for 8x8 computations.

Include/arm_nnfunctions.h:1523

arm_cmsis_nn_status arm_convolve_1x1_out_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Optimised convolution for 1x1 output images (shape of BX1x1xC_OUT) for 8x8 computations.

  • Supported framework : TensorFlow Lite Micro

  • Optimised for Bx1×1xC output CNN layers.

  • Constraints:

    1. output_dims->h and output_dims->w must equal 1
    2. output_dims->c is expected to be a multiple of 4 for best performance
Parameters of arm_convolve_1x1_out_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that supplies a scratch buffer for activation rearrangement. A NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. The buffer must hold one 4-byte-aligned GEMM row, that is round_up_4(filter_dims->h * filter_dims->w * filter_dims->c) bytes, as returned by `arm_convolve_1x1_out_s8_get_buffer_size()`. The requirement does not scale with the group count: the kernel rewinds its im2col cursor to the start of the buffer after each group. Setting ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check, which is what TFLite Micro and derivatives do today. The caller is expected to clear the buffer, if applicable, for security reasons.
weight_sum_ctxconst cmsis_nn_context *inPer-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. The contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR; on other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (stride, dilation, pad, offsets). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantisation multipliers and shifts.
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inPointer to input data. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, KH, KW, C_IN]
filter_dataconst int8_t *inPointer to filter data. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, 1, 1, C_OUT]
output_dataint8_t *outPointer to output data. Data type: int8
Returns of arm_convolve_1x1_out_s8
Description
`ARM_CMSIS_NN_ARG_ERROR` on bad args, or `ARM_CMSIS_NN_SUCCESS` on success.
function

Get the required scratch buffer size for armconvolve1x1outs8().

Include/arm_nnfunctions.h:1554

int32_t arm_convolve_1x1_out_s8_get_buffer_size(const cmsis_nn_dims *filter_dims)

Get the required scratch buffer size for arm_convolve_1x1_out_s8().

Parameters of arm_convolve_1x1_out_s8_get_buffer_size
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, KH, KW, C_IN]
Returns of arm_convolve_1x1_out_s8_get_buffer_size
Description
For valid (non-negative, in-range) filter dimensions, the buffer size in bytes: round_up_4(KH * KW * C_IN) on builds with the MVE extension (ARM_MATH_MVEI), 0 otherwise, since `arm_convolve_1x1_out_s8()` only exists on MVE builds. Returns -1 if any of filter_dims->w, filter_dims->h or filter_dims->c is negative or out of int32_t range, or if the rounded-up product exceeds INT32_MAX. The validation runs on every build target, not just the MVE leg, so the contract does not vary by target.
function

1xn convolution for s4 weights

Include/arm_nnfunctions.h:1592

arm_cmsis_nn_status arm_convolve_1_x_n_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

1xn convolution for s4 weights

  • Supported framework : TensorFlow Lite Micro

  • The following constrains on the arguments apply

    1. stride.w * input_dims->c is a multiple of 4
    2. Explicit constraints(since it is for 1xN convolution) -## input_dims->h equals 1 -## output_dims->h equals 1 -## filter_dims->h equals 1
Parameters of arm_convolve_1_x_n_s4
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. arm_convolve_1_x_n_s4_get_buffer_size will return the buffer_size if required The caller is expected to clear the buffer, if applicable, for security reasons.
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension
filter_dataconst int8_t *inFilter data pointer. Data type: int8 as packed int4
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inOptional bias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_convolve_1_x_n_s4
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Get the required additional buffer size for 1xn convolution.

Include/arm_nnfunctions.h:1621

int32_t arm_convolve_1_x_n_s8_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required additional buffer size for 1xn convolution.

Parameters of arm_convolve_1_x_n_s8_get_buffer_size
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_1_x_n_s8_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or conv_params->stride.w is not positive. On builds with the MVE extension (ARM_MATH_MVEI) that is the staging size of `arm_convolve_1_x_n_s8()`, at least filter W * C_IN bytes, or -1 if it would not fit in an int32_t; other builds return `arm_convolve_s8_get_buffer_size()`.
function

Get the required additional buffer size for 1xn convolution.

Include/arm_nnfunctions.h:1643

int32_t arm_convolve_1_x_n_s4_get_buffer_size(
const cmsis_nn_conv_params *conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get the required additional buffer size for 1xn convolution.

Parameters of arm_convolve_1_x_n_s4_get_buffer_size
NameTypeDirectionDescription
conv_paramsconst cmsis_nn_conv_params *inConvolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_convolve_1_x_n_s4_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or conv_params->stride.w is not positive. It also returns -1 if the required size would not fit in an int32_t; on a Helium build the route whose padding lines up with the stride needs no buffer and returns 0 without computing one.
function

Wrapper function to pick the right optimized s8 depthwise convolution function.

Include/arm_nnfunctions.h:1731

arm_cmsis_nn_status arm_depthwise_conv_wrapper_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Wrapper function to pick the right optimized s8 depthwise convolution function.

  • Supported framework: TensorFlow Lite

  • Picks one of the the following functions

    1. arm_depthwise_conv_s8()
    2. arm_depthwise_conv_3x3_s8() - Cortex-M CPUs with DSP extension only
    3. arm_depthwise_conv_s8_opt()
  • Check details of arm_depthwise_conv_s8_opt() for potential data that can be accessed outside of the boundary.

Parameters of arm_depthwise_conv_wrapper_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s8_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s8_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer, if applicable, for security reasons.
weight_sum_ctxconst cmsis_nn_context *inPer-channel weight sums, supplied by the caller. The selected kernel only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter, bias and dw_conv_params->input_offset are unchanged - see `arm_depthwise_convolve_weight_sum()` for the layout and the full reuse rules. Whether the buffer is consumed at all depends on the route this wrapper takes. It is forwarded to `arm_depthwise_conv_s8_opt()`, which reads it under MVE, only when dw_conv_params->ch_mult == 1, input_dims->n == 1, and either both dilations are 1 or the layer is 1D and dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, dw_conv_params->padding.h == 0, dilation.h == 1 and dilation.w >= 1. Such a dilated 1D layer therefore reads the sums too. Outside those cases the wrapper calls `arm_depthwise_conv_s8()`, which has no such parameter and ignores the context entirely - which is why several in-tree tests legitimately pass sums built by `arm_convolve_weight_sum()`, or none at all, on those routes (a 2D-dilated layer, for example). On MVE with input_dims->c == 1 and an output channel count above CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (8 on armclang, 1 otherwise), the layer is instead converted to a regular convolution, and conv-style sums from `arm_convolve_weight_sum()` are what that route wants. Where the sums are actually read, fill the buffer with `arm_depthwise_convolve_weight_sum()`, passing dw_conv_params->input_offset as lhs_offset and the same bias given here, so that entry j holds input_offset * sum(weights of channel j) + bias[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. On the `arm_depthwise_conv_s8_opt()` route, a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on builds where the buffer is actually read (ARM_MATH_DSP and ARM_MATH_MVEI both defined); on other builds the parameter is unread and NULL is accepted. On MVE with input_dims->c == 1 and an output channel count above CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (8 on armclang, 1 otherwise), this wrapper instead diverts to `arm_convolve_wrapper_s8()`. That diversion exists only on MVE, and every kernel it can dispatch to diagnoses a NULL buf with ARM_CMSIS_NN_ARG_ERROR, so that route is covered too. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions; see weight_sum_ctx for which dilated layers take the `arm_depthwise_conv_s8_opt()` route. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_wrapper_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` on successful completion, or `ARM_CMSIS_NN_ARG_ERROR` on the `arm_depthwise_conv_s8_opt()` route if ctx->buf is NULL when a scratch buffer is required, or if weight_sum_ctx->buf is NULL on builds where it is read (ARM_MATH_DSP and ARM_MATH_MVEI both defined), or if ctx->size is non-zero and below `arm_depthwise_conv_s8_opt_get_buffer_size()` for a layer its channel path runs, or if that sizer returns -1 (a negative dimension or a byte count it cannot represent), or on the MVE `arm_convolve_wrapper_s8()` diversion route if weight_sum_ctx->buf is NULL.
function

Wrapper function to pick the right optimized s4 depthwise convolution function.

Include/arm_nnfunctions.h:1779

arm_cmsis_nn_status arm_depthwise_conv_wrapper_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Wrapper function to pick the right optimized s4 depthwise convolution function.

  • Supported framework: TensorFlow Lite
Parameters of arm_depthwise_conv_wrapper_s4
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s4_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s4_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer ,if applicable, for security reasons.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43].
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_wrapper_s4
Description
The function returns `ARM_CMSIS_NN_SUCCESS` - Successful completion.
function

Get size of additional buffer required by armdepthwiseconvwrappers8().

Include/arm_nnfunctions.h:1818

int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get size of additional buffer required by arm_depthwise_conv_wrapper_s8().

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_depthwise_conv_wrapper_s8_get_buffer_size
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_wrapper_s8_get_buffer_size
Description
Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid.
function

Get size of additional buffer required by armdepthwiseconvwrappers8() for processors with DSP extension.

Include/arm_nnfunctions.h:1834

int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get size of additional buffer required by arm_depthwise_conv_wrapper_s8() for processors with DSP extension.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp
Description
Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid.
function

Get size of additional buffer required by armdepthwiseconvwrappers8() for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:1850

int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size_mve(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get size of additional buffer required by arm_depthwise_conv_wrapper_s8() for Arm(R) Helium Architecture case.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_depthwise_conv_wrapper_s8_get_buffer_size_mve
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_wrapper_s8_get_buffer_size_mve
Description
Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid.
function

Get size of additional buffer required by armdepthwiseconvwrappers4().

Include/arm_nnfunctions.h:1878

int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get size of additional buffer required by arm_depthwise_conv_wrapper_s4().

This sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as arm_depthwise_conv_s8_opt_get_buffer_size(), so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_depthwise_conv_wrapper_s4_get_buffer_size
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_wrapper_s4_get_buffer_size
Description
Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid.
function

Get size of additional buffer required by armdepthwiseconvwrappers4() for processors with DSP extension.

Include/arm_nnfunctions.h:1895

int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get size of additional buffer required by arm_depthwise_conv_wrapper_s4() for processors with DSP extension.

This sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as arm_depthwise_conv_s8_opt_get_buffer_size(), so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp
Description
Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid.
function

Get size of additional buffer required by armdepthwiseconvwrappers4() for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:1913

int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size_mve(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get size of additional buffer required by arm_depthwise_conv_wrapper_s4() for Arm(R) Helium Architecture case.

This sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as arm_depthwise_conv_s8_opt_get_buffer_size(), so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_depthwise_conv_wrapper_s4_get_buffer_size_mve
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_wrapper_s4_get_buffer_size_mve
Description
Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid.
function

Basic s8 depthwise convolution function that doesn't have any constraints on the input dimensions.

Include/arm_nnfunctions.h:1947

arm_cmsis_nn_status arm_depthwise_conv_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Basic s8 depthwise convolution function that doesn’t have any constraints on the input dimensions.

  • Supported framework: TensorFlow Lite
Parameters of arm_depthwise_conv_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s8_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s8()`, size the context with `arm_depthwise_conv_wrapper_s8_get_buffer_size()` instead, because another route through that wrapper does require a buffer.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Basic s4 depthwise convolution function that doesn't have any constraints on the input dimensions.

Include/arm_nnfunctions.h:1989

arm_cmsis_nn_status arm_depthwise_conv_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input,
const cmsis_nn_dims *filter_dims,
const int8_t *kernel,
const cmsis_nn_dims *bias_dims,
const int32_t *bias,
const cmsis_nn_dims *output_dims,
int8_t *output
)

Basic s4 depthwise convolution function that doesn’t have any constraints on the input dimensions.

  • Supported framework: TensorFlow Lite
Parameters of arm_depthwise_conv_s4
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s4_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s4()`, size the context with `arm_depthwise_conv_wrapper_s4_get_buffer_size()` instead, because another route through that wrapper does require a buffer.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.
inputconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
kernelconst int8_t *inFilter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43].
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
biasconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
outputint8_t *in, outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_s4
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Basic s16 depthwise convolution function that doesn't have any constraints on the input dimensions.

Include/arm_nnfunctions.h:2029

arm_cmsis_nn_status arm_depthwise_conv_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

Basic s16 depthwise convolution function that doesn’t have any constraints on the input dimensions.

  • Supported framework: TensorFlow Lite
Parameters of arm_depthwise_conv_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s16_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s16()`, size the context with `arm_depthwise_conv_wrapper_s16_get_buffer_size()` instead, because another route through that wrapper does require a buffer.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) conv_params->input_offset : Not used conv_params->output_offset : Not used
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.
input_dataconst int16_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int64_t *inBias data pointer. Data type: int64
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_depthwise_conv_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Wrapper function to pick the right optimized s16 depthwise convolution function.

Include/arm_nnfunctions.h:2082

arm_cmsis_nn_status arm_depthwise_conv_wrapper_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

Wrapper function to pick the right optimized s16 depthwise convolution function.

  • Supported framework: TensorFlow Lite

  • Picks one of the the following functions

    1. arm_depthwise_conv_s16()
    2. arm_depthwise_conv_fast_s16() - Cortex-M CPUs with DSP extension only
Parameters of arm_depthwise_conv_wrapper_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s16_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s16_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer, if applicable, for security reasons.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions. When ch_mult == 1 and filter_dims->w * filter_dims->h < 512, `arm_depthwise_conv_fast_s16()` is used for an undilated layer and for a 1D layer dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, dw_conv_params->padding.h == 0, dilation.h == 1 and dilation.w >= 1. Other layers use `arm_depthwise_conv_s16()`. Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->output_offset : Not used
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int64_t *inBias data pointer. Data type: int64
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_depthwise_conv_wrapper_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS` - Successful completion.
function

Get size of additional buffer required by armdepthwiseconvwrappers16().

Include/arm_nnfunctions.h:2119

int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get size of additional buffer required by arm_depthwise_conv_wrapper_s16().

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_depthwise_conv_wrapper_s16_get_buffer_size
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_wrapper_s16_get_buffer_size
Description
Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid.
function

Get size of additional buffer required by armdepthwiseconvwrappers16() for processors with DSP extension.

Include/arm_nnfunctions.h:2135

int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get size of additional buffer required by arm_depthwise_conv_wrapper_s16() for processors with DSP extension.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp
Description
Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid.
function

Get size of additional buffer required by armdepthwiseconvwrappers16() for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:2152

int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size_mve(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Get size of additional buffer required by arm_depthwise_conv_wrapper_s16() for Arm(R) Helium Architecture case.

Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.

Parameters of arm_depthwise_conv_wrapper_s16_get_buffer_size_mve
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_wrapper_s16_get_buffer_size_mve
Description
Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid.
function

Optimized s16 depthwise convolution function with constraint that inchannel equals outchannel.

Include/arm_nnfunctions.h:2200

arm_cmsis_nn_status arm_depthwise_conv_fast_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

Optimized s16 depthwise convolution function with constraint that in_channel equals out_channel.

ARM_CMSIS_NN_SUCCESS - Successful operation

  • Supported framework: TensorFlow Lite

  • The following constraints on the arguments apply

    1. ch_mult == 1: the number of input channels equals the number of output channels
    2. filter_dims->w * filter_dims->h < MAX_COL_COUNT (512)
    3. dw_conv_params->dilation.h == 1 and dw_conv_params->dilation.w >= 1
  • Recommended when number of channels is 4 or greater.

Parameters of arm_depthwise_conv_fast_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. `arm_depthwise_conv_fast_s16_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation.w is honoured; dw_conv_params->dilation.h must be 1. dw_conv_params->input_offset : Not used dw_conv_params->output_offset : Not used
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int64_t *inBias data pointer. Data type: int64
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_depthwise_conv_fast_s16
Description
The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - ctx-buff == NULL and `arm_depthwise_conv_fast_s16_get_buffer_size()` != 0 or input channel != output channel or filter_dims->w * filter_dims->h >= MAX_COL_COUNT (512) or dw_conv_params->dilation.h != 1 or dw_conv_params->dilation.w < 1
function

Get the required buffer size for optimized s16 depthwise convolution function with constraint that inchannel equals outchannel.

Include/arm_nnfunctions.h:2225

int32_t arm_depthwise_conv_fast_s16_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for optimized s16 depthwise convolution function with constraint that in_channel equals out_channel.

The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.

Parameters of arm_depthwise_conv_fast_s16_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_fast_s16_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below).

Include/arm_nnfunctions.h:2261

arm_cmsis_nn_status arm_depthwise_conv_3x3_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below).

  • Supported framework : TensorFlow Lite Micro

  • The following constrains on the arguments apply

    1. Number of input channel equals number of output channels
    2. Filter height and width equals 3
    3. Padding along x is either 0 or 1.
Parameters of arm_depthwise_conv_3x3_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. This kernel uses no additional buffer, so ctx->buf may be NULL.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_3x3_s8
Description
The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - Unsupported dimension of tensors - Unsupported pad size along the x axis `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

Optimized s8 depthwise convolution function with constraint that inchannel equals outchannel.

Include/arm_nnfunctions.h:2347

arm_cmsis_nn_status arm_depthwise_conv_s8_opt(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Optimized s8 depthwise convolution function with constraint that in_channel equals out_channel.

  • Supported framework: TensorFlow Lite

  • The following constrains on the arguments apply

    1. Number of input channel equals number of output channels or ch_mult equals 1
  • Reccomended when number of channels is 4 or greater.

  • On builds with ARM_MATH_DSP and ARM_MATH_MVEI, layers that arm_depthwise_conv_s8_opt_planar_supported() accepts run the planar path, with the same result as arm_depthwise_conv_s8_opt_planar(), unless ctx->size cannot hold its plane; every other layer runs the channel path of arm_depthwise_conv_s8_opt_channelwise(). Callers that choose the path ahead of time can call either one directly.

Parameters of arm_depthwise_conv_s8_opt
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer if required by the function. `arm_depthwise_conv_s8_opt_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.
weight_sum_ctxconst cmsis_nn_context *inPer-channel weight sums, supplied by the caller and only read by this function. See the note below for how to size, fill and reuse the buffer and for when a NULL buf is diagnosed.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation.w is honoured; dw_conv_params->dilation.h must be 1. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_s8_opt
Description
The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - input channel != output channel or ch_mult != 1, or dw_conv_params->dilation.h != 1 or dw_conv_params->dilation.w < 1, or ctx->buf is NULL when a scratch buffer is required, or ctx->size is non-zero and below `arm_depthwise_conv_s8_opt_get_buffer_size()` for a layer the channel path runs, or that sizer returns -1 (a negative dimension or a byte count it cannot represent) on the channel path, or weight_sum_ctx->buf is NULL on builds where it is read (ARM_MATH_DSP and ARM_MATH_MVEI both defined) `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

Whether armdepthwiseconvs8opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels.

Include/arm_nnfunctions.h:2384

int32_t arm_depthwise_conv_s8_opt_planar_supported(
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims
)

Whether arm_depthwise_conv_s8_opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels.

  • The rule is plain C and evaluates the same on every build, so a code generator can apply it ahead of time. The planar path itself exists only on builds with ARM_MATH_DSP and ARM_MATH_MVEI.
  • It depends only on the shapes and dw_conv_params: batch 1, C_IN equal to C_OUT, positive dimensions, stride 1, ch_mult 1, dilation.h 1, dilation.w at least 1 and at most 128 / C (integer division), at most 32 channels, the widths the path is faster for, and a plane that fits the scratch. This function is the reference for the rule; a mirror should be checked against it.
  • The width thresholds follow measured speed and may be retuned in a later release. A caller that calls arm_depthwise_conv_s8_opt_planar() directly must handle ARM_CMSIS_NN_NO_IMPL_ERROR, for example by calling arm_depthwise_conv_s8_opt_channelwise().
Parameters of arm_depthwise_conv_s8_opt_planar_supported
NameTypeDirectionDescription
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
Returns of arm_depthwise_conv_s8_opt_planar_supported
Description
1 when the planar path takes the layer with a scratch of `arm_depthwise_conv_s8_opt_get_buffer_size_mve()` bytes, 0 otherwise.
function

The planar path of armdepthwiseconvs8opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane.

Include/arm_nnfunctions.h:2420

arm_cmsis_nn_status arm_depthwise_conv_s8_opt_planar(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

The planar path of arm_depthwise_conv_s8_opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane.

  • The output is identical to arm_depthwise_conv_s8_opt() and arm_depthwise_conv_s8(). The bias is read through the weight sums; bias_dims and bias_data are unused.
Parameters of arm_depthwise_conv_s8_opt_planar
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context with the `arm_depthwise_conv_s8_opt_get_buffer_size()` scratch
weight_sum_ctxconst cmsis_nn_context *inPer-channel weight sums, as for `arm_depthwise_conv_s8_opt()`
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()`
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_s8_opt_planar
Description
The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_depthwise_conv_s8_opt()`, except its channel-path ctx->size check: a ctx->size too small for the plane returns ARM_CMSIS_NN_NO_IMPL_ERROR instead `ARM_CMSIS_NN_NO_IMPL_ERROR` - `arm_depthwise_conv_s8_opt_planar_supported()` rejects the layer, ctx->size cannot hold its plane, or the build lacks ARM_MATH_DSP or ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

The channel-vectorized path of armdepthwiseconvs8opt() on its own, without the planar attempt.

Include/arm_nnfunctions.h:2457

arm_cmsis_nn_status arm_depthwise_conv_s8_opt_channelwise(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

The channel-vectorized path of arm_depthwise_conv_s8_opt() on its own, without the planar attempt.

  • The output is identical to arm_depthwise_conv_s8_opt() and arm_depthwise_conv_s8() for every layer.
Parameters of arm_depthwise_conv_s8_opt_channelwise
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context with the `arm_depthwise_conv_s8_opt_get_buffer_size()` scratch
weight_sum_ctxconst cmsis_nn_context *inPer-channel weight sums, as for `arm_depthwise_conv_s8_opt()`
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()`
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_s8_opt_channelwise
Description
The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_depthwise_conv_s8_opt()` `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of armdepthwiseconvs8opt(), for layers in its gate.

Include/arm_nnfunctions.h:2513

arm_cmsis_nn_status arm_depthwise_conv_s8_opt_3x3(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of arm_depthwise_conv_s8_opt(), for layers in its gate. It computes three output rows per weight load.

  • The output is identical to arm_depthwise_conv_s8_opt() and arm_depthwise_conv_s8(). The bias is read through the weight sums, which arm_depthwise_convolve_weight_sum() fills as for arm_depthwise_conv_s8_opt(); bias_dims and bias_data are unused.
  • Gate: filter 3x3, dilation 1, N 1, C_IN equal to C_OUT with 16 <= C <= 2048 and C % 4 == 0, stride 1 or 2 and padding 0 or 1 in each dimension, input W >= 3 and H >= 1, output H >= 3 and W x H >= 16, every dimension at most 4096 and each tensor at most INT32_MAX elements, and the centre of the last output column’s window inside the input: (output W - 1) x stride.w - padding.w + 1 < input W. Input rows above or below the input count as padding, as in arm_depthwise_conv_s8().
  • Buffers: ctx and weight_sum_ctx and their buf are non-NULL, and ctx->size is at least arm_depthwise_conv_s8_opt_3x3_get_buffer_size() (3008 + input W x C + 16 bytes). ctx->buf needs no alignment.
  • It is a direct entry: arm_depthwise_conv_s8_opt() does not call it. A caller that selects the kernel per layer ahead of time calls arm_depthwise_conv_s8_opt_3x3_c64_s1() for C 64 with stride.h 1, this function for the rest of the gate, and arm_depthwise_conv_s8_opt() for every other layer, or on ARM_CMSIS_NN_NO_IMPL_ERROR. All three take the same arguments and the same weight sums; a ctx that serves all three holds the larger of arm_depthwise_conv_s8_opt_3x3_get_buffer_size() and arm_depthwise_conv_s8_opt_get_buffer_size(). On ARM_MATH_MVEI builds the second is the larger for a 3x3 filter when input W x C <= 1440.
Parameters of arm_depthwise_conv_s8_opt_3x3
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context with at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` bytes of scratch
weight_sum_ctxconst cmsis_nn_context *inPer-channel weight sums, as for `arm_depthwise_conv_s8_opt()`
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()`
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_s8_opt_3x3
Description
The function returns one of the following `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer or a buffer is outside the gate below, or the build lacks ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

armdepthwiseconvs8opt3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps.

Include/arm_nnfunctions.h:2558

arm_cmsis_nn_status arm_depthwise_conv_s8_opt_3x3_c64_s1(
const cmsis_nn_context *ctx,
const cmsis_nn_context *weight_sum_ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

arm_depthwise_conv_s8_opt_3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps. It is the faster entry for those layers.

  • The output is identical to arm_depthwise_conv_s8_opt_3x3(), arm_depthwise_conv_s8_opt() and arm_depthwise_conv_s8(). The bias is read through the weight sums; bias_dims and bias_data are unused.
  • The two entries share only their gate, parameter packing and the code for the output_y % 3 remainder rows, so a build with -ffunction-sections and section garbage collection keeps only the code of the entries it calls.
Parameters of arm_depthwise_conv_s8_opt_3x3_c64_s1
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context with at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` bytes of scratch
weight_sum_ctxconst cmsis_nn_context *inPer-channel weight sums, as for `arm_depthwise_conv_s8_opt()`
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()`
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_s8_opt_3x3_c64_s1
Description
The function returns one of the following `ARM_CMSIS_NN_NO_IMPL_ERROR` - C_IN is not 64, dw_conv_params->stride.h is not 1, the layer or a buffer is outside the gate of `arm_depthwise_conv_s8_opt_3x3()`, or the build lacks ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

Get the scratch size in bytes of armdepthwiseconvs8opt3x3() and armdepthwiseconvs8opt3x3c64s1().

Include/arm_nnfunctions.h:2585

int32_t arm_depthwise_conv_s8_opt_3x3_get_buffer_size(const cmsis_nn_dims *input_dims)

Get the scratch size in bytes of arm_depthwise_conv_s8_opt_3x3() and arm_depthwise_conv_s8_opt_3x3_c64_s1().

  • The size is the minimum ctx->size both entries accept. It depends only on the input width and channel count, since the filter is always 3x3.
  • The function is plain C and returns the same size on every build, including builds without ARM_MATH_MVEI where the entries return ARM_CMSIS_NN_NO_IMPL_ERROR, so a code generator can size the scratch ahead of time. A non-negative size is not a statement that the layer is in the gate of the entries.
Parameters of arm_depthwise_conv_s8_opt_3x3_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]. Only W and C_IN are read.
Returns of arm_depthwise_conv_s8_opt_3x3_get_buffer_size
Description
3008 + W x C_IN + 16 bytes, or -1 if W or C_IN is negative or the size would not fit in an int32_t
function

Optimized s4 depthwise convolution function with constraint that inchannel equals outchannel.

Include/arm_nnfunctions.h:2626

arm_cmsis_nn_status arm_depthwise_conv_s4_opt(
const cmsis_nn_context *ctx,
const cmsis_nn_dw_conv_params *dw_conv_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Optimized s4 depthwise convolution function with constraint that in_channel equals out_channel.

  • Supported framework: TensorFlow Lite

  • The following constrains on the arguments apply

    1. Number of input channel equals number of output channels or ch_mult equals 1
  • Reccomended when number of channels is 4 or greater.

Parameters of arm_depthwise_conv_s4_opt
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context that contains the additional buffer required by the function. `arm_depthwise_conv_s4_opt_get_buffer_size()` will return the buffer_size. A NULL ctx->buf is diagnosed with `ARM_CMSIS_NN_ARG_ERROR`. The caller is expected to clear the buffer, if applicable, for security reasons.
dw_conv_paramsconst cmsis_nn_dw_conv_params *inDepthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
filter_dataconst int8_t *inFilter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43].
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT]
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_depthwise_conv_s4_opt
Description
The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - input channel != output channel or ch_mult != 1 `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel.

Include/arm_nnfunctions.h:2651

int32_t arm_depthwise_conv_s8_opt_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel.

The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.

Parameters of arm_depthwise_conv_s8_opt_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_s8_opt_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t
function

Get the required buffer size for optimized s4 depthwise convolution function with constraint that inchannel equals outchannel.

Include/arm_nnfunctions.h:2667

int32_t arm_depthwise_conv_s4_opt_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)

Get the required buffer size for optimized s4 depthwise convolution function with constraint that in_channel equals out_channel.

The dimensions are not checked here: the query routes straight to the s8 _mve/_dsp leg and relies on the range checks inside that leg. Both legs apply the same check as arm_depthwise_conv_s8_opt_get_buffer_size(), so the answer for an out-of-range shape is the same on every build target.

Parameters of arm_depthwise_conv_s4_opt_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [1, H, W, C_OUT]
Returns of arm_depthwise_conv_s4_opt_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if input_dims->c or a filter dimension it reads is negative, or the required size would not fit in an int32_t.
function

Basic s4 Fully Connected function.

Include/arm_nnfunctions.h:2717

arm_cmsis_nn_status arm_fully_connected_s4(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Basic s4 Fully Connected function.

  • Supported framework: TensorFlow Lite
Parameters of arm_fully_connected_s4
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_fully_connected_s4_get_buffer_size(). Do not size this context with `arm_fully_connected_s8_get_buffer_size()`: that sizes the kernel-sum buffer of a different kernel and does not describe this argument.
fc_paramsconst cmsis_nn_fc_params *inFully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_tensor_quant_params *inPer-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor.
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inTwo dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used
filter_dataconst int8_t *inFilter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43].
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT] N, H, W : Not used
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_fully_connected_s4
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Basic s8 Fully Connected function.

Include/arm_nnfunctions.h:2789

arm_cmsis_nn_status arm_fully_connected_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Basic s8 Fully Connected function.

  • Supported framework: TensorFlow Lite
Parameters of arm_fully_connected_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inPer-output-channel kernel sums, supplied by the caller - not scratch memory that this function fills in. The library never populates ctx->buf for this function, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds this function adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. Sized by `arm_fully_connected_s8_get_buffer_size()`: filter_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call.
fc_paramsconst cmsis_nn_fc_params *inFully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_tensor_quant_params *inPer-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor.
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inTwo dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT] N, H, W : Not used
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_fully_connected_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Basic s8 Fully Connected function using per channel quantization.

Include/arm_nnfunctions.h:2862

arm_cmsis_nn_status arm_fully_connected_per_channel_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

Basic s8 Fully Connected function using per channel quantization.

  • Supported framework: TensorFlow Lite
Parameters of arm_fully_connected_per_channel_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inPer-output-channel kernel sums, supplied by the caller - not scratch memory that this function fills in. The library never populates ctx->buf for this function, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds this function adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. There is no per-channel sizer; `arm_fully_connected_s8_get_buffer_size()` returns the same quantity this function needs, filter_dims->c * sizeof(int32_t) where the sums are used and 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call.
fc_paramsconst cmsis_nn_fc_params *inFully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inTwo dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT] N, H, W : Not used
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_fully_connected_per_channel_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

s8 Fully Connected layer wrapper function

Include/arm_nnfunctions.h:2937

arm_cmsis_nn_status arm_fully_connected_wrapper_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 Fully Connected layer wrapper function

  • Supported framework: TensorFlow Lite
Parameters of arm_fully_connected_wrapper_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inPer-output-channel kernel sums, supplied by the caller - not scratch memory that this wrapper fills in. The library never populates ctx->buf here, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. The context is passed straight through to `arm_fully_connected_per_channel_s8()` or `arm_fully_connected_s8()` depending on quant_params->is_per_channel, and both read it the same way. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. Neither selected kernel writes the buffer, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds the selected kernel adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. There is no wrapper sizer; `arm_fully_connected_s8_get_buffer_size()` returns the quantity both routes need, filter_dims->c * sizeof(int32_t) where the sums are used and 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call.
fc_paramsconst cmsis_nn_fc_params *inFully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127]
quant_paramsconst cmsis_nn_quant_params *inPer-channel or per-tensor quantization info. Check struct defintion for details. It contains the multiplier and shift value(s) to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inTwo dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT] N, H, W : Not used
bias_dataconst int32_t *inBias data pointer. Data type: int32
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_fully_connected_wrapper_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s32 biasdata.

Include/arm_nnfunctions.h:2961

arm_cmsis_nn_status arm_vector_sum_s8(
int32_t *vector_sum_buf,
const int32_t vector_cols,
const int32_t vector_rows,
const int8_t *vector_data,
const int32_t lhs_offset,
const int32_t rhs_offset,
const int32_t *bias_data
)

Calculate the sum of each row in vector_data, multiply by lhs_offset and optionally add s32 bias_data.

Parameters of arm_vector_sum_s8
NameTypeDirectionDescription
vector_sum_bufint32_t *in, outBuffer for vector sums
vector_colsconst int32_tinNumber of vector columns
vector_rowsconst int32_tinNumber of vector rows
vector_dataconst int8_t *inVector of weigths data
lhs_offsetconst int32_tinConstant multiplied with each sum
rhs_offsetconst int32_tinConstant added to each vector element before sum
bias_dataconst int32_t *inVector of bias data, added to each sum.
Returns of arm_vector_sum_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s64 biasdata.

Include/arm_nnfunctions.h:2980

arm_cmsis_nn_status arm_vector_sum_s8_s64(
int64_t *vector_sum_buf,
const int32_t vector_cols,
const int32_t vector_rows,
const int8_t *vector_data,
const int32_t lhs_offset,
const int64_t *bias_data
)

Calculate the sum of each row in vector_data, multiply by lhs_offset and optionally add s64 bias_data.

Parameters of arm_vector_sum_s8_s64
NameTypeDirectionDescription
vector_sum_bufint64_t *in, outBuffer for vector sums
vector_colsconst int32_tinNumber of vector columns
vector_rowsconst int32_tinNumber of vector rows
vector_dataconst int8_t *inVector of weigths data
lhs_offsetconst int32_tinConstant multiplied with each sum
bias_dataconst int64_t *inVector of bias data, added to each sum.
Returns of arm_vector_sum_s8_s64
Description
The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

Get size of additional buffer required by armfullyconnecteds8().

Include/arm_nnfunctions.h:2998

int32_t arm_fully_connected_s8_get_buffer_size(const cmsis_nn_dims *filter_dims)

Get size of additional buffer required by arm_fully_connected_s8(). See also arm_vector_sum_s8, which is required if buffer size is > 0.

For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.

Parameters of arm_fully_connected_s8_get_buffer_size
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *indimension of filter
Returns of arm_fully_connected_s8_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t
function

Get size of additional buffer required by armfullyconnecteds8() for processors with DSP extension.

Include/arm_nnfunctions.h:3009

int32_t arm_fully_connected_s8_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)

Get size of additional buffer required by arm_fully_connected_s8() for processors with DSP extension.

For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.

Parameters of arm_fully_connected_s8_get_buffer_size_dsp
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *indimension of filter
Returns of arm_fully_connected_s8_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t
function

Get size of additional buffer required by armfullyconnecteds8() for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:3020

int32_t arm_fully_connected_s8_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)

Get size of additional buffer required by arm_fully_connected_s8() for Arm(R) Helium Architecture case.

For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.

Parameters of arm_fully_connected_s8_get_buffer_size_mve
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *indimension of filter
Returns of arm_fully_connected_s8_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t
function

Basic s16 Fully Connected function.

Include/arm_nnfunctions.h:3057

arm_cmsis_nn_status arm_fully_connected_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

Basic s16 Fully Connected function.

  • Supported framework: TensorFlow Lite
Parameters of arm_fully_connected_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused. This function currently ignores the context entirely on every build - it neither reads nor writes ctx->buf - and `arm_fully_connected_s16_get_buffer_size()` returns 0 accordingly, so { NULL, 0 } is accepted. Unlike the s8 variants, no precomputed kernel sums are required here. None of this is a guarantee about future versions.
fc_paramsconst cmsis_nn_fc_params *inFully Connected layer parameters. fc_params->input_offset : 0 fc_params->filter_offset : 0 fc_params->output_offset : 0
quant_paramsconst cmsis_nn_per_tensor_quant_params *inPer-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor.
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inTwo dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT] N, H, W : Not used
bias_dataconst int64_t *inBias data pointer. Data type: int64
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_fully_connected_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Basic s16 Fully Connected function using per channel quantization.

Include/arm_nnfunctions.h:3114

arm_cmsis_nn_status arm_fully_connected_per_channel_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_per_channel_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *kernel,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

Basic s16 Fully Connected function using per channel quantization.

  • Supported framework: TensorFlow Lite
Parameters of arm_fully_connected_per_channel_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outScratch buffer that this function writes before it reads, on every build. It is filled here with one reduced int32 multiplier per output channel derived from quant_params->multiplier, so the caller supplies the storage only and the incoming contents are never used. Unlike the s8 variants, no precomputed kernel sums are expected, and clearing the buffer is harmless. Required on every build, not only under MVE: ctx or ctx->buf being NULL is reported as ARM_CMSIS_NN_ARG_ERROR, as is a non-zero ctx->size smaller than the requirement. A ctx->size of 0 is treated as undeclared and is not checked. Sized by `arm_fully_connected_per_channel_s16_get_buffer_size()`: filter_dims->c * sizeof(int32_t), which equals the output_dims->c entries written. The caller is expected to clear the buffer afterwards, if applicable, for security reasons.
fc_paramsconst cmsis_nn_fc_params *inFully Connected layer parameters. Range of fc_params->input_offset : 0 fc_params->filter_offset : 0 Range of fc_params->output_offset : 0
quant_paramsconst cmsis_nn_per_channel_quant_params *inPer-channel quantization info. It contains the multiplier and shift values to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inTwo dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used
kernelconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT] N, H, W : Not used
bias_dataconst int64_t *inBias data pointer. Data type: int64
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_fully_connected_per_channel_s16
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

s16 Fully Connected layer wrapper function

Include/arm_nnfunctions.h:3174

arm_cmsis_nn_status arm_fully_connected_wrapper_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_fc_params *fc_params,
const cmsis_nn_quant_params *quant_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const int8_t *filter_data,
const cmsis_nn_dims *bias_dims,
const int64_t *bias_data,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

s16 Fully Connected layer wrapper function

  • Supported framework: TensorFlow Lite
Parameters of arm_fully_connected_wrapper_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outScratch buffer, whose use depends on the route taken. Unlike the s8 wrapper, no precomputed kernel sums are expected on either route, and clearing the buffer is harmless. When quant_params->is_per_channel is set, the context is passed to `arm_fully_connected_per_channel_s16()`, which writes it before reading it, on every build: it is filled there with one reduced int32 multiplier per output channel, so the caller supplies the storage only. On that route ctx or ctx->buf being NULL is reported as ARM_CMSIS_NN_ARG_ERROR, as is a non-zero ctx->size smaller than filter_dims->c * sizeof(int32_t); a ctx->size of 0 is treated as undeclared and is not checked. Size it with `arm_fully_connected_per_channel_s16_get_buffer_size()`. Otherwise the context goes to `arm_fully_connected_s16()`, which currently ignores it entirely, so { NULL, 0 } is accepted on that route. A caller that does not know the route in advance should size for the per-channel case, since `arm_fully_connected_s16_get_buffer_size()` returns 0. None of this is a guarantee about future versions. The caller is expected to clear the buffer afterwards, if applicable, for security reasons.
fc_paramsconst cmsis_nn_fc_params *inFully Connected layer parameters. Range of fc_params->input_offset : 0 fc_params->filter_offset : 0 Range of fc_params->output_offset : 0
quant_paramsconst cmsis_nn_quant_params *inPer-channel or per-tensor quantization info. Check struct defintion for details. It contains the multiplier and shift value(s) to be applied to each output channel
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inTwo dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used
filter_dataconst int8_t *inFilter data pointer. Data type: int8
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions. Format: [C_OUT] N, H, W : Not used
bias_dataconst int64_t *inBias data pointer. Data type: int64
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_fully_connected_wrapper_s16
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Get size of additional buffer required by armfullyconnecteds16().

Include/arm_nnfunctions.h:3192

int32_t arm_fully_connected_s16_get_buffer_size(const cmsis_nn_dims *filter_dims)

Get size of additional buffer required by arm_fully_connected_s16().

Parameters of arm_fully_connected_s16_get_buffer_size
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *indimension of filter
Returns of arm_fully_connected_s16_get_buffer_size
Description
The function returns required buffer size in bytes
function

Get size of additional buffer required by armfullyconnecteds16() for processors with DSP extension.

Include/arm_nnfunctions.h:3202

int32_t arm_fully_connected_s16_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)

Get size of additional buffer required by arm_fully_connected_s16() for processors with DSP extension.

Parameters of arm_fully_connected_s16_get_buffer_size_dsp
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *indimension of filter
Returns of arm_fully_connected_s16_get_buffer_size_dsp
Description
The function returns required buffer size in bytes
function

Get size of additional buffer required by armfullyconnecteds16() for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:3212

int32_t arm_fully_connected_s16_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)

Get size of additional buffer required by arm_fully_connected_s16() for Arm(R) Helium Architecture case.

Parameters of arm_fully_connected_s16_get_buffer_size_mve
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *indimension of filter
Returns of arm_fully_connected_s16_get_buffer_size_mve
Description
The function returns required buffer size in bytes
function

Get size of additional buffer required by armfullyconnectedperchannels16().

Include/arm_nnfunctions.h:3223

int32_t arm_fully_connected_per_channel_s16_get_buffer_size(const cmsis_nn_dims *filter_dims)

Get size of additional buffer required by arm_fully_connected_per_channel_s16().

For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.

Parameters of arm_fully_connected_per_channel_s16_get_buffer_size
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *indimension of filter
Returns of arm_fully_connected_per_channel_s16_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t
function

Get size of additional buffer required by armfullyconnectedperchannels16() for processors with DSP extension.

Include/arm_nnfunctions.h:3235

int32_t arm_fully_connected_per_channel_s16_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)

Get size of additional buffer required by arm_fully_connected_per_channel_s16() for processors with DSP extension.

For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.

Parameters of arm_fully_connected_per_channel_s16_get_buffer_size_dsp
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *indimension of filter
Returns of arm_fully_connected_per_channel_s16_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t
function

Get size of additional buffer required by armfullyconnectedperchannels16() for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:3247

int32_t arm_fully_connected_per_channel_s16_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)

Get size of additional buffer required by arm_fully_connected_per_channel_s16() for Arm(R) Helium Architecture case.

For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.

Parameters of arm_fully_connected_per_channel_s16_get_buffer_size_mve
NameTypeDirectionDescription
filter_dimsconst cmsis_nn_dims *indimension of filter
Returns of arm_fully_connected_per_channel_s16_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t
function

s8 elementwise add of two tensors with support for broadcasting.

Include/arm_nnfunctions.h:3285

arm_cmsis_nn_status arm_add_s8(
const int8_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int8_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)

s8 elementwise add of two tensors with support for broadcasting.

Parameters of arm_add_s8
NameTypeDirectionDescription
input1_dataconst int8_t *inpointer to input tensor 1
input1_dimsconst cmsis_nn_dims *inpointer to input tensor 1 dimensions
input2_dataconst int8_t *inpointer to input tensor 2
input2_dimsconst cmsis_nn_dims *inpointer to input tensor 2 dimensions
input1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input1_multconst int32_tinmultiplier for input 1
input1_shiftconst int32_tinshift for input 1
input2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input2_multconst int32_tinmultiplier for input 2
input2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinleft shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
output_dataint8_t *outpointer to output tensor
output_dimsconst cmsis_nn_dims *inpointer to output tensor dimensions
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
Returns of arm_add_s8
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape.
function

s8 elementwise add of scalar and vector

Include/arm_nnfunctions.h:3329

arm_cmsis_nn_status arm_add_scalar_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s8 elementwise add of scalar and vector

Parameters of arm_add_scalar_s8
NameTypeDirectionDescription
input_1_vectconst int8_t *inpointer to input scalar
input_2_vectconst int8_t *inpointer to input vector
input_1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinleft shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint8_t *outpointer to output vector
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
block_sizeconst int32_tinnumber of samples
Returns of arm_add_scalar_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s8 elementwise add of two vectors

Include/arm_nnfunctions.h:3369

arm_cmsis_nn_status arm_elementwise_add_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s8 elementwise add of two vectors

Parameters of arm_elementwise_add_s8
NameTypeDirectionDescription
input_1_vectconst int8_t *inpointer to input vector 1
input_2_vectconst int8_t *inpointer to input vector 2
input_1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tininput left shift. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint8_t *in, outpointer to output vector
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
block_sizeconst int32_tinnumber of samples
Returns of arm_elementwise_add_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s8 elementwise absolute value

Include/arm_nnfunctions.h:3400

arm_cmsis_nn_status arm_abs_s8(
const int8_t *input,
const int32_t input_offset,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const bool needs_rescale,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s8 elementwise absolute value

Parameters of arm_abs_s8
NameTypeDirectionDescription
inputconst int8_t *inpointer to input vector
input_offsetconst int32_tininput offset
outputint8_t *outpointer to output vector
out_offsetconst int32_tinoutput offset
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
needs_rescaleconst boolinindicates if output requantization is needed
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
block_sizeconst int32_tinnumber of samples
Returns of arm_abs_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s8 elementwise square root

Include/arm_nnfunctions.h:3420

arm_cmsis_nn_status arm_sqrt_s8(
const int8_t *input,
const cmsis_nn_dims *input_dims,
int8_t *output,
const int8_t *sqrt_lut
)

s8 elementwise square root

Parameters of arm_sqrt_s8
NameTypeDirectionDescription
inputconst int8_t *inpointer to input vector
input_dimsconst cmsis_nn_dims *inpointer to input tensor dimensions
outputint8_t *outpointer to output vector
sqrt_lutconst int8_t *inpointer to 256-entry lookup table
Returns of arm_sqrt_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise square root using piecewise LUT with linear interpolation

Include/arm_nnfunctions.h:3431

arm_cmsis_nn_status arm_sqrt_s16(
const int16_t *input,
const cmsis_nn_dims *input_dims,
int16_t *output,
const int16_t *sqrt_lut
)

s16 elementwise square root using piecewise LUT with linear interpolation

Parameters of arm_sqrt_s16
NameTypeDirectionDescription
inputconst int16_t *inpointer to input vector
input_dimsconst cmsis_nn_dims *inpointer to input tensor dimensions
outputint16_t *outpointer to output vector
sqrt_lutconst int16_t *inpointer to 513-entry lookup table (int16_t)
Returns of arm_sqrt_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise square root without a lookup table

Include/arm_nnfunctions.h:3454

arm_cmsis_nn_status arm_sqrt_s16_tablefree(
const int16_t *input,
const cmsis_nn_dims *input_dims,
int16_t *output,
const float scale
)

s16 elementwise square root without a lookup table

Approximates output[i] = trunc(sqrt(input[i] * scale)) saturated to 32767, which is LiteRT’s int16 SQRT (dequantize in float32, sqrtf, divide by the output scale, truncate, clamp) for zero points 0, to within 1 LSB of LiteRT at every non-negative input for input scales 1e-7 to 1e-1 and output scales from 0.01x to 10x the full-range scale, saturating ones included. Inputs at or below 0 produce

  1. Needs no table; the int16 API does not depend on ARM_NN_ENABLE_F32/F16, and on targets without a floating-point unit the plain C path uses fmaf from the C library.
Parameters of arm_sqrt_s16_tablefree
NameTypeDirectionDescription
inputconst int16_t *inpointer to input vector
input_dimsconst cmsis_nn_dims *inpointer to input tensor dimensions
outputint16_t *outpointer to output vector
scaleconst floatininput_scale / (output_scale * output_scale) as float32: take the float32-rounded tensor scales, evaluate in float64 and round once to float32. Must be finite and greater than 0.
Returns of arm_sqrt_s16_tablefree
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise absolute value

Include/arm_nnfunctions.h:3470

arm_cmsis_nn_status arm_abs_s16(
const int16_t *input,
const int32_t input_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const bool needs_rescale,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise absolute value

Parameters of arm_abs_s16
NameTypeDirectionDescription
inputconst int16_t *inpointer to input vector
input_offsetconst int32_tininput offset
outputint16_t *outpointer to output vector
out_offsetconst int32_tinoutput offset
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
needs_rescaleconst boolinindicates if output requantization is needed
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_abs_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

INT16 reciprocal square root using a per-operator LUT.

Include/arm_nnfunctions.h:3496

arm_cmsis_nn_status arm_rsqrt_s16_per_op(
const int16_t *input,
const int32_t input_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size,
const int16_t *lut
)

INT16 reciprocal square root using a per-operator LUT.

Parameters of arm_rsqrt_s16_per_op
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input buffer.
input_offsetconst int32_tinInput tensor zero offset. The kernel evaluates each element as `input - input_offset` before the LUT lookup.
outputint16_t *outPointer to the output buffer.
out_offsetconst int32_tinOutput tensor zero offset.
out_activation_minconst int32_tinMinimum output clamp.
out_activation_maxconst int32_tinMaximum output clamp.
block_sizeconst int32_tinNumber of elements.
lutconst int16_t *inPointer to a 513-entry INT16 LUT in output domain.
Returns of arm_rsqrt_s16_per_op
Description
The function returns ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR.
function

INT16 reciprocal square root using a shared universal LUT.

Include/arm_nnfunctions.h:3530

arm_cmsis_nn_status arm_rsqrt_s16_universal(
const int16_t *input,
const int32_t input_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const bool needs_rescale,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size,
const int32_t *lut
)

INT16 reciprocal square root using a shared universal LUT.

In universal mode all RSQRT operators share a single LUT that captures the base 1/sqrt(x) shape, and operator-specific quantization is applied afterward via out_mult / out_shift. Because this two-step process introduces extra rounding stages, the output may differ from the per-op variant (arm_rsqrt_s16_per_op) by up to ±3 LSB per element. This is expected and acceptable for deployment.

Parameters of arm_rsqrt_s16_universal
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input buffer.
input_offsetconst int32_tinInput tensor zero offset. The kernel evaluates each element as `input - input_offset` before the LUT lookup.
outputint16_t *outPointer to the output buffer.
out_offsetconst int32_tinOutput tensor zero offset.
out_multconst int32_tinOutput requantization multiplier.
out_shiftconst int32_tinOutput requantization shift.
needs_rescaleconst boolinWhether requantization is required.
out_activation_minconst int32_tinMinimum output clamp.
out_activation_maxconst int32_tinMaximum output clamp.
block_sizeconst int32_tinNumber of elements.
lutconst int32_t *inPointer to a 513-entry INT32 shared LUT in Q30 domain.
Returns of arm_rsqrt_s16_universal
Description
The function returns ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR.
function

s8 elementwise subtraction of two tensors with support for broadcasting.

Include/arm_nnfunctions.h:3571

arm_cmsis_nn_status arm_sub_s8(
const int8_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int8_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)

s8 elementwise subtraction of two tensors with support for broadcasting.

Parameters of arm_sub_s8
NameTypeDirectionDescription
input1_dataconst int8_t *inpointer to input tensor 1
input1_dimsconst cmsis_nn_dims *inpointer to input tensor 1 dimensions
input2_dataconst int8_t *inpointer to input tensor 2
input2_dimsconst cmsis_nn_dims *inpointer to input tensor 2 dimensions
input1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input1_multconst int32_tinmultiplier for input 1
input1_shiftconst int32_tinshift for input 1
input2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input2_multconst int32_tinmultiplier for input 2
input2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinleft shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
output_dataint8_t *outpointer to output tensor
output_dimsconst cmsis_nn_dims *inpointer to output tensor dimensions
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
Returns of arm_sub_s8
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape.
function

s8 elementwise subtract of scalar and vector (scalar - vector)

Include/arm_nnfunctions.h:3615

arm_cmsis_nn_status arm_sub_scalar_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s8 elementwise subtract of scalar and vector (scalar - vector)

Parameters of arm_sub_scalar_s8
NameTypeDirectionDescription
input_1_vectconst int8_t *inpointer to input scalar
input_2_vectconst int8_t *inpointer to input vector
input_1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinleft shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint8_t *outpointer to output vector
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
block_sizeconst int32_tinnumber of samples
Returns of arm_sub_scalar_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s8 elementwise subtract of two vectors

Include/arm_nnfunctions.h:3655

arm_cmsis_nn_status arm_elementwise_sub_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s8 elementwise subtract of two vectors

Parameters of arm_elementwise_sub_s8
NameTypeDirectionDescription
input_1_vectconst int8_t *inpointer to input vector 1
input_2_vectconst int8_t *inpointer to input vector 2
input_1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tininput left shift. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint8_t *in, outpointer to output vector
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
block_sizeconst int32_tinnumber of samples
Returns of arm_elementwise_sub_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise add of two tensors with support for broadcasting.

Include/arm_nnfunctions.h:3702

arm_cmsis_nn_status arm_add_s16(
const int16_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int16_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)

s16 elementwise add of two tensors with support for broadcasting.

Parameters of arm_add_s16
NameTypeDirectionDescription
input1_dataconst int16_t *inpointer to input tensor 1
input1_dimsconst cmsis_nn_dims *inpointer to input tensor 1 dimensions
input2_dataconst int16_t *inpointer to input tensor 2
input2_dimsconst cmsis_nn_dims *inpointer to input tensor 2 dimensions
input1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input1_multconst int32_tinmultiplier for input 1
input1_shiftconst int32_tinshift for input 1
input2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input2_multconst int32_tinmultiplier for input 2
input2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinleft shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
output_dataint16_t *outpointer to output tensor
output_dimsconst cmsis_nn_dims *inpointer to output tensor dimensions
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
Returns of arm_add_s16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape.
function

s16 elementwise add of scalar and vector

Include/arm_nnfunctions.h:3747

arm_cmsis_nn_status arm_add_scalar_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise add of scalar and vector

Parameters of arm_add_scalar_s16
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input scalar
input_2_vectconst int16_t *inpointer to input vector
input_1_offsetconst int32_tinoffset for input 1. Not used.
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Not used.
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinleft shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint16_t *outpointer to output vector
out_offsetconst int32_tinoutput offset. Not used.
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_add_scalar_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise add of two vectors

Include/arm_nnfunctions.h:3789

arm_cmsis_nn_status arm_elementwise_add_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise add of two vectors

Parameters of arm_elementwise_add_s16
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input vector 1
input_2_vectconst int16_t *inpointer to input vector 2
input_1_offsetconst int32_tinoffset for input 1. Not used.
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Not used.
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tininput left shift. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint16_t *in, outpointer to output vector
out_offsetconst int32_tinoutput offset. Not used.
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_elementwise_add_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise subtraction of two tensors with support for broadcasting.

Include/arm_nnfunctions.h:3836

arm_cmsis_nn_status arm_sub_s16(
const int16_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int16_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)

s16 elementwise subtraction of two tensors with support for broadcasting.

Parameters of arm_sub_s16
NameTypeDirectionDescription
input1_dataconst int16_t *inpointer to input tensor 1
input1_dimsconst cmsis_nn_dims *inpointer to input tensor 1 dimensions
input2_dataconst int16_t *inpointer to input tensor 2
input2_dimsconst cmsis_nn_dims *inpointer to input tensor 2 dimensions
input1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input1_multconst int32_tinmultiplier for input 1
input1_shiftconst int32_tinshift for input 1
input2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input2_multconst int32_tinmultiplier for input 2
input2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinleft shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
output_dataint16_t *outpointer to output tensor
output_dimsconst cmsis_nn_dims *inpointer to output tensor dimensions
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
Returns of arm_sub_s16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape.
function

s16 elementwise subtract of scalar and vector (scalar - vector)

Include/arm_nnfunctions.h:3881

arm_cmsis_nn_status arm_sub_scalar_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise subtract of scalar and vector (scalar - vector)

Parameters of arm_sub_scalar_s16
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input scalar
input_2_vectconst int16_t *inpointer to input vector
input_1_offsetconst int32_tinoffset for input 1. Not used.
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Not used.
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinleft shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint16_t *outpointer to output vector
out_offsetconst int32_tinoutput offset. Not used.
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_sub_scalar_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise subtract of two vectors

Include/arm_nnfunctions.h:3923

arm_cmsis_nn_status arm_elementwise_sub_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise subtract of two vectors

Parameters of arm_elementwise_sub_s16
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input vector 1
input_2_vectconst int16_t *inpointer to input vector 2
input_1_offsetconst int32_tinoffset for input 1. Not used.
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Not used.
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tininput left shift. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint16_t *in, outpointer to output vector
out_offsetconst int32_tinoutput offset. Not used.
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_elementwise_sub_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s8 elementwise squared difference of two tensors with support for broadcasting.

Include/arm_nnfunctions.h:3969

arm_cmsis_nn_status arm_squared_difference_s8(
const int8_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int8_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)

s8 elementwise squared difference of two tensors with support for broadcasting.

Parameters of arm_squared_difference_s8
NameTypeDirectionDescription
input1_dataconst int8_t *inpointer to input tensor 1
input1_dimsconst cmsis_nn_dims *inpointer to input tensor 1 dimensions
input2_dataconst int8_t *inpointer to input tensor 2
input2_dimsconst cmsis_nn_dims *inpointer to input tensor 2 dimensions
input1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input1_multconst int32_tinmultiplier for input 1
input1_shiftconst int32_tinshift for input 1
input2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input2_multconst int32_tinmultiplier for input 2
input2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinCommon left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
output_dataint8_t *outpointer to output tensor
output_dimsconst cmsis_nn_dims *inpointer to output tensor dimensions
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
Returns of arm_squared_difference_s8
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape.
function

s8 elementwise squared difference of scalar and vector.

Include/arm_nnfunctions.h:4013

arm_cmsis_nn_status arm_squared_difference_scalar_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s8 elementwise squared difference of scalar and vector.

Parameters of arm_squared_difference_scalar_s8
NameTypeDirectionDescription
input_1_vectconst int8_t *inpointer to input scalar
input_2_vectconst int8_t *inpointer to input vector
input_1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinCommon left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint8_t *outpointer to output vector
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
block_sizeconst int32_tinnumber of samples
Returns of arm_squared_difference_scalar_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s8 elementwise squared difference of two vectors.

Include/arm_nnfunctions.h:4055

arm_cmsis_nn_status arm_elementwise_squared_difference_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s8 elementwise squared difference of two vectors.

Parameters of arm_elementwise_squared_difference_s8
NameTypeDirectionDescription
input_1_vectconst int8_t *inpointer to input vector 1
input_2_vectconst int8_t *inpointer to input vector 2
input_1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinCommon left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint8_t *outpointer to output vector
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
block_sizeconst int32_tinnumber of samples
Returns of arm_elementwise_squared_difference_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise squared difference of two tensors with support for broadcasting.

Include/arm_nnfunctions.h:4101

arm_cmsis_nn_status arm_squared_difference_s16(
const int16_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int16_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input1_mult,
const int32_t input1_shift,
const int32_t input2_offset,
const int32_t input2_mult,
const int32_t input2_shift,
const int32_t left_shift,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)

s16 elementwise squared difference of two tensors with support for broadcasting.

Parameters of arm_squared_difference_s16
NameTypeDirectionDescription
input1_dataconst int16_t *inpointer to input tensor 1
input1_dimsconst cmsis_nn_dims *inpointer to input tensor 1 dimensions
input2_dataconst int16_t *inpointer to input tensor 2
input2_dimsconst cmsis_nn_dims *inpointer to input tensor 2 dimensions
input1_offsetconst int32_tinoffset for input 1
input1_multconst int32_tinmultiplier for input 1
input1_shiftconst int32_tinshift for input 1
input2_offsetconst int32_tinoffset for input 2
input2_multconst int32_tinmultiplier for input 2
input2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinCommon left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
output_dataint16_t *outpointer to output tensor
output_dimsconst cmsis_nn_dims *inpointer to output tensor dimensions
out_offsetconst int32_tinoutput offset
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
Returns of arm_squared_difference_s16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape.
function

s16 elementwise squared difference of scalar and vector.

Include/arm_nnfunctions.h:4145

arm_cmsis_nn_status arm_squared_difference_scalar_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise squared difference of scalar and vector.

Parameters of arm_squared_difference_scalar_s16
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input scalar
input_2_vectconst int16_t *inpointer to input vector
input_1_offsetconst int32_tinoffset for input 1
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinCommon left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint16_t *outpointer to output vector
out_offsetconst int32_tinoutput offset
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_squared_difference_scalar_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise squared difference of two vectors.

Include/arm_nnfunctions.h:4187

arm_cmsis_nn_status arm_elementwise_squared_difference_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise squared difference of two vectors.

Parameters of arm_elementwise_squared_difference_s16
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input vector 1
input_2_vectconst int16_t *inpointer to input vector 2
input_1_offsetconst int32_tinoffset for input 1
input_1_multconst int32_tinmultiplier for input 1
input_1_shiftconst int32_tinshift for input 1
input_2_offsetconst int32_tinoffset for input 2
input_2_multconst int32_tinmultiplier for input 2
input_2_shiftconst int32_tinshift for input 2
left_shiftconst int32_tinCommon left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
outputint16_t *outpointer to output vector
out_offsetconst int32_tinoutput offset
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_elementwise_squared_difference_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s8 elementwise multiplication of two tensors with support for broadcasting.

Include/arm_nnfunctions.h:4224

arm_cmsis_nn_status arm_mul_s8(
const int8_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int8_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input2_offset,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)

s8 elementwise multiplication of two tensors with support for broadcasting.

Parameters of arm_mul_s8
NameTypeDirectionDescription
input1_dataconst int8_t *inpointer to input tensor 1
input1_dimsconst cmsis_nn_dims *inpointer to input tensor 1 dimensions
input2_dataconst int8_t *inpointer to input tensor 2
input2_dimsconst cmsis_nn_dims *inpointer to input tensor 2 dimensions
input1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
output_dataint8_t *outpointer to output tensor
output_dimsconst cmsis_nn_dims *inpointer to output tensor dimensions
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
Returns of arm_mul_s8
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape.
function

s8 elementwise multiplication of scalar and vector

Include/arm_nnfunctions.h:4254

arm_cmsis_nn_status arm_mul_scalar_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s8 elementwise multiplication of scalar and vector

Parameters of arm_mul_scalar_s8
NameTypeDirectionDescription
input_1_vectconst int8_t *inpointer to input scalar
input_2_vectconst int8_t *inpointer to input vector
input_1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input_2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
outputint8_t *outpointer to output vector
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
block_sizeconst int32_tinnumber of samples
Returns of arm_mul_scalar_s8
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s8 elementwise multiplication

Include/arm_nnfunctions.h:4283

arm_cmsis_nn_status arm_elementwise_mul_s8(
const int8_t *input_1_vect,
const int8_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int8_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s8 elementwise multiplication

Supported framework: TensorFlow Lite micro

Parameters of arm_elementwise_mul_s8
NameTypeDirectionDescription
input_1_vectconst int8_t *inpointer to input vector 1
input_2_vectconst int8_t *inpointer to input vector 2
input_1_offsetconst int32_tinoffset for input 1. Range: -127 to 128
input_2_offsetconst int32_tinoffset for input 2. Range: -127 to 128
outputint8_t *in, outpointer to output vector
out_offsetconst int32_tinoutput offset. Range: -128 to 127
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -128
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 127
block_sizeconst int32_tinnumber of samples
Returns of arm_elementwise_mul_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s16 elementwise multiplication of two tensors with support for broadcasting.

Include/arm_nnfunctions.h:4315

arm_cmsis_nn_status arm_mul_s16(
const int16_t *input1_data,
const cmsis_nn_dims *input1_dims,
const int16_t *input2_data,
const cmsis_nn_dims *input2_dims,
const int32_t input1_offset,
const int32_t input2_offset,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max
)

s16 elementwise multiplication of two tensors with support for broadcasting.

Parameters of arm_mul_s16
NameTypeDirectionDescription
input1_dataconst int16_t *inpointer to input tensor 1
input1_dimsconst cmsis_nn_dims *inpointer to input tensor 1 dimensions
input2_dataconst int16_t *inpointer to input tensor 2
input2_dimsconst cmsis_nn_dims *inpointer to input tensor 2 dimensions
input1_offsetconst int32_tinoffset for input 1. Not used.
input2_offsetconst int32_tinoffset for input 2. Not used.
output_dataint16_t *outpointer to output tensor
output_dimsconst cmsis_nn_dims *inpointer to output tensor dimensions
out_offsetconst int32_tinoutput offset. Not used.
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
Returns of arm_mul_s16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape.
function

s16 elementwise multiplication of scalar and vector

Include/arm_nnfunctions.h:4345

arm_cmsis_nn_status arm_mul_scalar_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise multiplication of scalar and vector

Parameters of arm_mul_scalar_s16
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input scalar
input_2_vectconst int16_t *inpointer to input vector
input_1_offsetconst int32_tinoffset for input 1. Not used.
input_2_offsetconst int32_tinoffset for input 2. Not used.
outputint16_t *outpointer to output vector
out_offsetconst int32_tinoutput offset. Not used.
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_mul_scalar_s16
Description
The function returns ARM_CMSIS_NN_SUCCESS
function

s16 elementwise multiplication

Include/arm_nnfunctions.h:4374

arm_cmsis_nn_status arm_elementwise_mul_s16(
const int16_t *input_1_vect,
const int16_t *input_2_vect,
const int32_t input_1_offset,
const int32_t input_2_offset,
int16_t *output,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift,
const int32_t out_activation_min,
const int32_t out_activation_max,
const int32_t block_size
)

s16 elementwise multiplication

Supported framework: TensorFlow Lite micro

Parameters of arm_elementwise_mul_s16
NameTypeDirectionDescription
input_1_vectconst int16_t *inpointer to input vector 1
input_2_vectconst int16_t *inpointer to input vector 2
input_1_offsetconst int32_tinoffset for input 1. Not used.
input_2_offsetconst int32_tinoffset for input 2. Not used.
outputint16_t *in, outpointer to output vector
out_offsetconst int32_tinoutput offset. Not used.
out_multconst int32_tinoutput multiplier
out_shiftconst int32_tinoutput shift
out_activation_minconst int32_tinminimum value to clamp output to. Min: -32768
out_activation_maxconst int32_tinmaximum value to clamp output to. Max: 32767
block_sizeconst int32_tinnumber of samples
Returns of arm_elementwise_mul_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

s8 elementwise minimum w/ support for broadcasting and scalar inputs.

Include/arm_nnfunctions.h:4405

arm_cmsis_nn_status arm_minimum_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)

s8 elementwise minimum w/ support for broadcasting and scalar inputs.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_minimum_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int8_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int8_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_dataint8_t *outPointer to the output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_minimum_s8
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

s8 elementwise maximum w/ support for broadcasting and scalar inputs.

Include/arm_nnfunctions.h:4432

arm_cmsis_nn_status arm_maximum_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)

s8 elementwise maximum w/ support for broadcasting and scalar inputs.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_maximum_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int8_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int8_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_dataint8_t *outPointer to the output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_maximum_s8
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

s16 elementwise minimum w/ support for broadcasting and scalar inputs.

Include/arm_nnfunctions.h:4459

arm_cmsis_nn_status arm_minimum_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)

s16 elementwise minimum w/ support for broadcasting and scalar inputs.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_minimum_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int16_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int16_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_dataint16_t *outPointer to the output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_minimum_s16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

s16 elementwise maximum w/ support for broadcasting and scalar inputs.

Include/arm_nnfunctions.h:4486

arm_cmsis_nn_status arm_maximum_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)

s16 elementwise maximum w/ support for broadcasting and scalar inputs.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_maximum_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int16_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int16_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_dataint16_t *outPointer to the output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_maximum_s16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

s8 elementwise comparison with support for broadcasting.

Include/arm_nnfunctions.h:4527

arm_cmsis_nn_status arm_comparison_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
arm_nn_compare_operation operation
)

s8 elementwise comparison with support for broadcasting.

Parameters of arm_comparison_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int8_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int8_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
operationarm_nn_compare_operationinComparison operation to perform
Returns of arm_comparison_s8
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

s16 elementwise comparison with support for broadcasting.

Include/arm_nnfunctions.h:4570

arm_cmsis_nn_status arm_comparison_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift,
arm_nn_compare_operation operation
)

s16 elementwise comparison with support for broadcasting.

Parameters of arm_comparison_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int16_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int16_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
operationarm_nn_compare_operationinComparison operation to perform
Returns of arm_comparison_s16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL.
function

s8 elementwise equality comparison with support for broadcasting.

Include/arm_nnfunctions.h:4610

arm_cmsis_nn_status arm_equal_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s8 elementwise equality comparison with support for broadcasting.

Parameters of arm_equal_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int8_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int8_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_equal_s8
Description
As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s8 elementwise inequality comparison with support for broadcasting.

Include/arm_nnfunctions.h:4649

arm_cmsis_nn_status arm_not_equal_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s8 elementwise inequality comparison with support for broadcasting.

Parameters of arm_not_equal_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int8_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int8_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_not_equal_s8
Description
As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s8 elementwise greater-than comparison with support for broadcasting.

Include/arm_nnfunctions.h:4688

arm_cmsis_nn_status arm_greater_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s8 elementwise greater-than comparison with support for broadcasting.

Parameters of arm_greater_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int8_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int8_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_greater_s8
Description
As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s8 elementwise greater-or-equal comparison with support for broadcasting.

Include/arm_nnfunctions.h:4727

arm_cmsis_nn_status arm_greater_equal_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s8 elementwise greater-or-equal comparison with support for broadcasting.

Parameters of arm_greater_equal_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int8_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int8_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_greater_equal_s8
Description
As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s8 elementwise less-than comparison with support for broadcasting.

Include/arm_nnfunctions.h:4766

arm_cmsis_nn_status arm_less_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s8 elementwise less-than comparison with support for broadcasting.

Parameters of arm_less_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int8_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int8_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_less_s8
Description
As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s8 elementwise less-or-equal comparison with support for broadcasting.

Include/arm_nnfunctions.h:4805

arm_cmsis_nn_status arm_less_equal_s8(
const cmsis_nn_context *ctx,
const int8_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int8_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s8 elementwise less-or-equal comparison with support for broadcasting.

Parameters of arm_less_equal_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int8_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int8_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_less_equal_s8
Description
As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s16 elementwise equality comparison with support for broadcasting.

Include/arm_nnfunctions.h:4844

arm_cmsis_nn_status arm_equal_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s16 elementwise equality comparison with support for broadcasting.

Parameters of arm_equal_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int16_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int16_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_equal_s16
Description
As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s16 elementwise inequality comparison with support for broadcasting.

Include/arm_nnfunctions.h:4883

arm_cmsis_nn_status arm_not_equal_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s16 elementwise inequality comparison with support for broadcasting.

Parameters of arm_not_equal_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int16_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int16_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_not_equal_s16
Description
As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s16 elementwise greater-than comparison with support for broadcasting.

Include/arm_nnfunctions.h:4922

arm_cmsis_nn_status arm_greater_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s16 elementwise greater-than comparison with support for broadcasting.

Parameters of arm_greater_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int16_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int16_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_greater_s16
Description
As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s16 elementwise greater-or-equal comparison with support for broadcasting.

Include/arm_nnfunctions.h:4961

arm_cmsis_nn_status arm_greater_equal_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s16 elementwise greater-or-equal comparison with support for broadcasting.

Parameters of arm_greater_equal_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int16_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int16_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_greater_equal_s16
Description
As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s16 elementwise less-than comparison with support for broadcasting.

Include/arm_nnfunctions.h:5000

arm_cmsis_nn_status arm_less_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s16 elementwise less-than comparison with support for broadcasting.

Parameters of arm_less_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int16_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int16_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_less_s16
Description
As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

s16 elementwise less-or-equal comparison with support for broadcasting.

Include/arm_nnfunctions.h:5039

arm_cmsis_nn_status arm_less_equal_s16(
const cmsis_nn_context *ctx,
const int16_t *input_1_data,
const cmsis_nn_dims *input_1_dims,
const int16_t *input_2_data,
const cmsis_nn_dims *input_2_dims,
bool *output_data,
const cmsis_nn_dims *output_dims,
const int32_t input_1_offset,
const int32_t input_1_mult,
const int32_t input_1_shift,
const int32_t input_2_offset,
const int32_t input_2_mult,
const int32_t input_2_shift,
const int32_t left_shift
)

s16 elementwise less-or-equal comparison with support for broadcasting.

Parameters of arm_less_equal_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused; may be NULL.
input_1_dataconst int16_t *inPointer to input1 tensor
input_1_dimsconst cmsis_nn_dims *inInput1 tensor dimensions
input_2_dataconst int16_t *inPointer to input2 tensor
input_2_dimsconst cmsis_nn_dims *inInput2 tensor dimensions
output_databool *outPointer to the output tensor (bool values)
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
input_1_offsetconst int32_tinZero-point for input1 tensor
input_1_multconst int32_tinMultiplier for input1 tensor
input_1_shiftconst int32_tinShift for input1 tensor
input_2_offsetconst int32_tinZero-point for input2 tensor
input_2_multconst int32_tinMultiplier for input2 tensor
input_2_shiftconst int32_tinShift for input2 tensor
left_shiftconst int32_tinCommon left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.
Returns of arm_less_equal_s16
Description
As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments.
function

Q7 RELU function.

Include/arm_nnfunctions.h:5067

void arm_relu_q7(int8_t *data, uint16_t size)

Q7 RELU function.

Parameters of arm_relu_q7
NameTypeDirectionDescription
dataint8_t *in, outpointer to input
sizeuint16_tinnumber of elements
function

Q7 RELU6 function.

Include/arm_nnfunctions.h:5074

void arm_relu6_q7(int8_t *data, uint16_t size)

Q7 RELU6 function.

Parameters of arm_relu6_q7
NameTypeDirectionDescription
dataint8_t *in, outpointer to input
sizeuint16_tinnumber of elements
function

Q15 RELU function.

Include/arm_nnfunctions.h:5081

void arm_relu_q15(int16_t *data, uint16_t size)

Q15 RELU function.

Parameters of arm_relu_q15
NameTypeDirectionDescription
dataint16_t *in, outpointer to input
sizeuint16_tinnumber of elements
function

S8 clamp function.

Include/arm_nnfunctions.h:5095

arm_cmsis_nn_status arm_clamp_s8(
const int8_t *input,
const int8_t act_min,
const int8_t act_max,
int8_t *output,
const int32_t output_size
)

S8 clamp function.

This function clamps each element in the input tensor to the range. This can be useful for activations such as relu(0, 6), relu(-1,1), etc.

Parameters of arm_clamp_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to input
act_minconst int8_tinMinimum value to clamp to
act_maxconst int8_tinMaximum value to clamp to
outputint8_t *outPointer to output
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_clamp_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

S16 clamp function.

Include/arm_nnfunctions.h:5113

arm_cmsis_nn_status arm_clamp_s16(
const int16_t *input,
const int16_t act_min,
const int16_t act_max,
int16_t *output,
const int32_t output_size
)

S16 clamp function.

This function clamps each element in the input tensor to the range. This can be useful for activations such as relu(0, 6), relu(-1,1), etc.

Parameters of arm_clamp_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to input
act_minconst int16_tinMinimum value to clamp to
act_maxconst int16_tinMaximum value to clamp to
outputint16_t *outPointer to output
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_clamp_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127.

Include/arm_nnfunctions.h:5132

arm_cmsis_nn_status arm_relu_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
int8_t *output,
const int32_t output_size
)

S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127.

Parameters of arm_relu_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input buffer
input_offsetconst int32_tinInput tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplierconst int32_tinOutput multiplier
output_shiftconst int32_tinOutput shift
outputint8_t *outPointer to the output buffer
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_relu_s8
Description
The function returns ARM_MATH_SUCCESS
function

S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.

Include/arm_nnfunctions.h:5155

arm_cmsis_nn_status arm_relu_generic_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
const int32_t act_min,
const int32_t act_max,
int8_t *output,
const int32_t output_size
)

S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.

Parameters of arm_relu_generic_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input buffer
input_offsetconst int32_tinInput tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplierconst int32_tinOutput multiplier
output_shiftconst int32_tinOutput shift
act_minconst int32_tinMinimum value to clamp the output to
act_maxconst int32_tinMaximum value to clamp the output to
outputint8_t *outPointer to the output buffer
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_relu_generic_s8
Description
The function returns ARM_MATH_SUCCESS
function

S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767.

Include/arm_nnfunctions.h:5178

arm_cmsis_nn_status arm_relu_s16(
const int16_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
int16_t *output,
const int32_t output_size
)

S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767.

Parameters of arm_relu_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input buffer
input_offsetconst int32_tinInput tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplierconst int32_tinOutput multiplier
output_shiftconst int32_tinOutput shift
outputint16_t *outPointer to the output buffer
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_relu_s16
Description
The function returns ARM_MATH_SUCCESS
function

S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.

Include/arm_nnfunctions.h:5201

arm_cmsis_nn_status arm_relu_generic_s16(
const int16_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
const int32_t act_min,
const int32_t act_max,
int16_t *output,
const int32_t output_size
)

S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.

Parameters of arm_relu_generic_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input buffer
input_offsetconst int32_tinInput tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplierconst int32_tinOutput multiplier
output_shiftconst int32_tinOutput shift
act_minconst int32_tinMinimum value to clamp the output to
act_maxconst int32_tinMaximum value to clamp the output to
outputint16_t *outPointer to the output buffer
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_relu_generic_s16
Description
The function returns ARM_MATH_SUCCESS
function

S8 Leaky ReLU activation function.

Include/arm_nnfunctions.h:5226

arm_cmsis_nn_status arm_leaky_relu_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
int8_t *output,
const int32_t output_size
)

S8 Leaky ReLU activation function.

Parameters of arm_leaky_relu_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input buffer
input_offsetconst int32_tinInput tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplier_alphaconst int32_tinOutput multiplier for the alpha parameter
output_shift_alphaconst int32_tinOutput shift for the alpha parameter
output_multiplier_identityconst int32_tinOutput multiplier for the identity parameter
output_shift_identityconst int32_tinOutput shift for the identity parameter
outputint8_t *outPointer to the output buffer
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_leaky_relu_s8
Description
The function returns ARM_MATH_SUCCESS
function

S16 Leaky ReLU activation function.

Include/arm_nnfunctions.h:5251

arm_cmsis_nn_status arm_leaky_relu_s16(
const int16_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
int16_t *output,
const int32_t output_size
)

S16 Leaky ReLU activation function.

Parameters of arm_leaky_relu_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input buffer
input_offsetconst int32_tinInput tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplier_alphaconst int32_tinOutput multiplier for the alpha parameter
output_shift_alphaconst int32_tinOutput shift for the alpha parameter
output_multiplier_identityconst int32_tinOutput multiplier for the identity parameter
output_shift_identityconst int32_tinOutput shift for the identity parameter
outputint16_t *outPointer to the output buffer
output_sizeconst int32_tinNumber of elements in the input tensor
Returns of arm_leaky_relu_s16
Description
The function returns ARM_MATH_SUCCESS
function

Logistic activation function for s16.

Include/arm_nnfunctions.h:5272

arm_cmsis_nn_status arm_logistic_s16(
const int16_t *input,
int16_t *output,
const int32_t input_size,
int32_t input_multiplier,
int32_t input_left_shift
)

Logistic activation function for s16.

Parameters of arm_logistic_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input tensor
outputint16_t *outPointer to the output tensor
input_sizeconst int32_tinNumber of elements in the input tensor
input_multiplierint32_tinInput quantization multiplier
input_left_shiftint32_tinInput quantization shift within the range [0, 31]
Returns of arm_logistic_s16
Description
The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

Tanh activation function for s16.

Include/arm_nnfunctions.h:5289

arm_cmsis_nn_status arm_tanh_s16(
const int16_t *input,
int16_t *output,
const int32_t input_size,
int32_t input_multiplier,
int32_t input_left_shift
)

Tanh activation function for s16.

Parameters of arm_tanh_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input tensor
outputint16_t *outPointer to the output tensor
input_sizeconst int32_tinNumber of elements in the input tensor
input_multiplierint32_tinInput quantization multiplier
input_left_shiftint32_tinInput quantization shift within the range [0, 31]
Returns of arm_tanh_s16
Description
The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

s16 neural network activation function using direct table look-up

Include/arm_nnfunctions.h:5308

arm_cmsis_nn_status arm_nn_activation_s16(
const int16_t *input,
int16_t *output,
const int32_t size,
const int32_t left_shift,
const arm_nn_activation_type type
)

s16 neural network activation function using direct table look-up

Supported framework: TensorFlow Lite for Microcontrollers. This activation function must be bit precise congruent with the corresponding TFLM tanh and sigmoid activation functions

Parameters of arm_nn_activation_s16
NameTypeDirectionDescription
inputconst int16_t *inpointer to input data
outputint16_t *outpointer to output
sizeconst int32_tinnumber of elements
left_shiftconst int32_tinbit-width of the integer part, assumed to be smaller than 3.
typeconst arm_nn_activation_typeintype of activation functions
Returns of arm_nn_activation_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

S8 Hard-Swish activation function (compatibility version).

Include/arm_nnfunctions.h:5338

arm_cmsis_nn_status arm_hard_swish_compat_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier_fp,
const int32_t output_multiplier_exp,
const int32_t relu_multiplier_fp,
const int32_t relu_multiplier_exp,
int8_t *output,
const int32_t output_size
)

S8 Hard-Swish activation function (compatibility version).

This version is compatible with TFLite implementation of Hard-Swish. hires_input_scale = (1.0 / 128.0) * float(input_scale) relu_scale = 3.0 / 32768.0 out_mul_real = hires_input_scale / float(output_scale) relu_mul_real = hires_input_scale / relu_scale output_multiplier_fp, output_multiplier_exp = to_q15_exp(out_mul_real) relu_multiplier_fp, relu_multiplier_exp = to_q15_exp(relu_mul_real) Here to_q15_exp quantizes to Q31 with a frexp exponent, then rounds and saturates the Q31 multiplier to Q15. For input_scale = output_scale = 0.125, the output pair is (16384, -6) and the ReLU pair is (21845, 4).

Parameters of arm_hard_swish_compat_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input buffer
input_offsetconst int32_tinInput tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplier_fpconst int32_tinOutput multiplier in fixed point format
output_multiplier_expconst int32_tinExponent for output multiplier
relu_multiplier_fpconst int32_tinReLU6 multiplier in fixed point format
relu_multiplier_expconst int32_tinExponent for ReLU6 multiplier
outputint8_t *outPointer to the output buffer
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_hard_swish_compat_s8
Description
ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR if output_multiplier_exp is positive.
function

S8 Hard-Swish activation function (precise version).

Include/arm_nnfunctions.h:5369

arm_cmsis_nn_status arm_hard_swish_precise_s8(
const int8_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
const int32_t relu_q3,
const int32_t relu_q6,
const int32_t prescale,
int8_t *output,
const int32_t output_size
)

S8 Hard-Swish activation function (precise version).

This version uses int32_t for intermediate computations to provide better accuracy. relu_q3, relu_q6 = round(3 / input_scale), round(6 / input_scale) prescale = min value to multiply input by to avoid overflow in intermediate computations M = ((input_scale**2) / (6.0 * output_scale)) * (1 << prescale) output_multiplier, output_shift = quantize_multiplier(M)

Parameters of arm_hard_swish_precise_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input buffer
input_offsetconst int32_tinInput tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplierconst int32_tinOutput multiplier
output_shiftconst int32_tinOutput shift
relu_q3const int32_tinReLU6 Q3 value
relu_q6const int32_tinReLU6 Q6 value
prescaleconst int32_tinPrescale to apply to input
outputint8_t *outPointer to the output buffer
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_hard_swish_precise_s8
Description
The function returns ARM_MATH_SUCCESS
function

S16 Hard-Swish activation function (precise version).

Include/arm_nnfunctions.h:5401

arm_cmsis_nn_status arm_hard_swish_precise_s16(
const int16_t *input,
const int32_t input_offset,
const int32_t output_offset,
const int32_t output_multiplier,
const int32_t output_shift,
const int32_t relu_q3,
const int32_t relu_q6,
const int32_t prescale,
int16_t *output,
const int32_t output_size
)

S16 Hard-Swish activation function (precise version).

This version uses int32_t for intermediate computations to provide better accuracy. relu_q3, relu_q6 = round(3 / input_scale), round(6 / input_scale) prescale = min value to multiply input by to avoid overflow in intermediate computations M = ((input_scale**2) / (6.0 * output_scale)) * (1 << prescale) output_multiplier, output_shift = quantize_multiplier(M)

Parameters of arm_hard_swish_precise_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input buffer
input_offsetconst int32_tinInput tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplierconst int32_tinOutput multiplier
output_shiftconst int32_tinOutput shift
relu_q3const int32_tinReLU6 Q3 value
relu_q6const int32_tinReLU6 Q6 value
prescaleconst int32_tinPrescale to apply to input
outputint16_t *outPointer to the output buffer
output_sizeconst int32_tinNumber of elements in the tensor
Returns of arm_hard_swish_precise_s16
Description
The function returns ARM_MATH_SUCCESS
function

S8 PReLU activation function.

Include/arm_nnfunctions.h:5432

arm_cmsis_nn_status arm_prelu_s8(
const cmsis_nn_dims *input_dims,
const int8_t *input,
const cmsis_nn_dims *alpha_dims,
const int8_t *alpha,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t output_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
const cmsis_nn_dims *output_dims,
int8_t *output
)

S8 PReLU activation function.

Parameters of arm_prelu_s8
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
inputconst int8_t *inPointer to the input buffer
alpha_dimsconst cmsis_nn_dims *inAlpha tensor dimensions. Format: [N, H, W, C]
alphaconst int8_t *inPointer to the alpha buffer
input_offsetconst int32_tinInput tensor zero offset
alpha_offsetconst int32_tinAlpha tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplier_identityconst int32_tinOutput multiplier 1
output_shift_identityconst int32_tinOutput shift 1
output_multiplier_alphaconst int32_tinOutput multiplier 2
output_shift_alphaconst int32_tinOutput shift 2
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
outputint8_t *outPointer to the output buffer
Returns of arm_prelu_s8
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, alpha does not broadcast into the input, or the output dimensions do not match the input dimensions.
function

Elementwise S8 PReLU activation function.

Include/arm_nnfunctions.h:5462

arm_cmsis_nn_status arm_elementwise_prelu_s8(
const int8_t *input,
const int8_t *alpha,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t out_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
int8_t *output,
const int32_t block_size
)

Elementwise S8 PReLU activation function.

Parameters of arm_elementwise_prelu_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input buffer
alphaconst int8_t *inPointer to the alpha buffer (same shape as input)
input_offsetconst int32_tinInput tensor zero offset
alpha_offsetconst int32_tinAlpha tensor zero offset
out_offsetconst int32_tinOutput tensor zero offset
output_multiplier_identityconst int32_tinOutput multiplier when input >= 0
output_shift_identityconst int32_tinOutput shift when input >= 0
output_multiplier_alphaconst int32_tinOutput multiplier when input < 0
output_shift_alphaconst int32_tinOutput shift when input < 0
outputint8_t *outPointer to the output buffer
block_sizeconst int32_tinNumber of elements to process
Returns of arm_elementwise_prelu_s8
Description
The function returns ARM_MATH_SUCCESS
function

Scalar S8 PReLU activation function.

Include/arm_nnfunctions.h:5491

arm_cmsis_nn_status arm_prelu_scalar_s8(
const int8_t *scalar_vect,
const int8_t *non_scalar_vect,
const bool scalar_is_input,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t output_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
int8_t *output,
const int32_t block_size
)

Scalar S8 PReLU activation function.

Parameters of arm_prelu_scalar_s8
NameTypeDirectionDescription
scalar_vectconst int8_t *inPointer to the scalar buffer (single value)
non_scalar_vectconst int8_t *inPointer to the non-scalar buffer
scalar_is_inputconst boolinTrue if the scalar buffer holds the input value, false if it holds alpha
input_offsetconst int32_tinInput tensor zero offset
alpha_offsetconst int32_tinAlpha tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplier_identityconst int32_tinOutput multiplier when input >= 0
output_shift_identityconst int32_tinOutput shift when input >= 0
output_multiplier_alphaconst int32_tinOutput multiplier when input < 0
output_shift_alphaconst int32_tinOutput shift when input < 0
outputint8_t *outPointer to the output buffer
block_sizeconst int32_tinNumber of elements to process when the non-scalar vector is used
Returns of arm_prelu_scalar_s8
Description
The function returns ARM_MATH_SUCCESS
function

S16 PReLU activation function.

Include/arm_nnfunctions.h:5524

arm_cmsis_nn_status arm_prelu_s16(
const cmsis_nn_dims *input_dims,
const int16_t *input,
const cmsis_nn_dims *alpha_dims,
const int16_t *alpha,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t output_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
const cmsis_nn_dims *output_dims,
int16_t *output
)

S16 PReLU activation function.

Parameters of arm_prelu_s16
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
inputconst int16_t *inPointer to the input buffer
alpha_dimsconst cmsis_nn_dims *inAlpha tensor dimensions. Format: [N, H, W, C]
alphaconst int16_t *inPointer to the alpha buffer
input_offsetconst int32_tinInput tensor zero offset
alpha_offsetconst int32_tinAlpha tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplier_identityconst int32_tinOutput multiplier when input >= 0
output_shift_identityconst int32_tinOutput shift when input >= 0
output_multiplier_alphaconst int32_tinOutput multiplier when input < 0
output_shift_alphaconst int32_tinOutput shift when input < 0
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C_OUT]
outputint16_t *outPointer to the output buffer
Returns of arm_prelu_s16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, alpha does not broadcast into the input, or the output dimensions do not match the input dimensions.
function

Elementwise S16 PReLU activation function.

Include/arm_nnfunctions.h:5554

arm_cmsis_nn_status arm_elementwise_prelu_s16(
const int16_t *input,
const int16_t *alpha,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t out_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
int16_t *output,
const int32_t block_size
)

Elementwise S16 PReLU activation function.

Parameters of arm_elementwise_prelu_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input buffer
alphaconst int16_t *inPointer to the alpha buffer (same shape as input)
input_offsetconst int32_tinInput tensor zero offset
alpha_offsetconst int32_tinAlpha tensor zero offset
out_offsetconst int32_tinOutput tensor zero offset
output_multiplier_identityconst int32_tinOutput multiplier when input >= 0
output_shift_identityconst int32_tinOutput shift when input >= 0
output_multiplier_alphaconst int32_tinOutput multiplier when input < 0
output_shift_alphaconst int32_tinOutput shift when input < 0
outputint16_t *outPointer to the output buffer
block_sizeconst int32_tinNumber of elements to process
Returns of arm_elementwise_prelu_s16
Description
ARM_CMSIS_NN_SUCCESS
function

Scalar S16 PReLU activation function.

Include/arm_nnfunctions.h:5583

arm_cmsis_nn_status arm_prelu_scalar_s16(
const int16_t *scalar_vect,
const int16_t *non_scalar_vect,
const bool scalar_is_input,
const int32_t input_offset,
const int32_t alpha_offset,
const int32_t output_offset,
const int32_t output_multiplier_identity,
const int32_t output_shift_identity,
const int32_t output_multiplier_alpha,
const int32_t output_shift_alpha,
int16_t *output,
const int32_t block_size
)

Scalar S16 PReLU activation function.

Parameters of arm_prelu_scalar_s16
NameTypeDirectionDescription
scalar_vectconst int16_t *inPointer to the scalar buffer (single value)
non_scalar_vectconst int16_t *inPointer to the non-scalar buffer
scalar_is_inputconst boolinTrue if the scalar buffer holds the input value, false if it holds alpha
input_offsetconst int32_tinInput tensor zero offset
alpha_offsetconst int32_tinAlpha tensor zero offset
output_offsetconst int32_tinOutput tensor zero offset
output_multiplier_identityconst int32_tinOutput multiplier when input >= 0
output_shift_identityconst int32_tinOutput shift when input >= 0
output_multiplier_alphaconst int32_tinOutput multiplier when input < 0
output_shift_alphaconst int32_tinOutput shift when input < 0
outputint16_t *outPointer to the output buffer
block_sizeconst int32_tinNumber of elements to process when the non-scalar vector is used
Returns of arm_prelu_scalar_s16
Description
ARM_CMSIS_NN_SUCCESS
function

s8 average pooling function.

Include/arm_nnfunctions.h:5638

arm_cmsis_nn_status arm_avgpool_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params *pool_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 average pooling function.

  • Supported Framework: TensorFlow Lite
Parameters of arm_avgpool_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context. Size ctx->buf with arm_avgpool_s8_get_buffer_size(output_dims->w, input_dims->c). Note that it takes the output width and the input channel count rather than a `cmsis_nn_dims`. `arm_avgpool_s8_get_buffer_size_dsp()` and `arm_avgpool_s8_get_buffer_size_mve()` size the same buffer for a specific target. It returns 0 where no buffer is needed, in which case ctx->buf may be NULL. The caller is expected to clear the buffer, if applicable, for security reasons.
pool_paramsconst cmsis_nn_pool_params *inPooling parameters
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [H, W] Argument N and C are not used.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN.
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_avgpool_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns (an extent of 0 or less), which writes nothing and does not use ctx `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments, including a negative channel count, a pooling window that does not overlap the input, window positions (output index times stride minus padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t, or, on builds without MVE, a NULL ctx, or a NULL ctx->buf where the sizer asks for a buffer. Nothing is written to output_data then.
function

Get the required buffer size for S8 average pooling function.

Include/arm_nnfunctions.h:5659

int32_t arm_avgpool_s8_get_buffer_size(const int dim_dst_width, const int ch_src)

Get the required buffer size for S8 average pooling function.

Unlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. arm_avgpool_s8() depends on that sentinel being non-zero, since it reads a non-zero size as “ctx->buf is required” before touching the accumulator buffer.

Parameters of arm_avgpool_s8_get_buffer_size
NameTypeDirectionDescription
dim_dst_widthconst intinoutput tensor dimension
ch_srcconst intinnumber of input tensor channels
Returns of arm_avgpool_s8_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t
function

Get the required buffer size for S8 average pooling function for processors with DSP extension.

Include/arm_nnfunctions.h:5671

int32_t arm_avgpool_s8_get_buffer_size_dsp(const int dim_dst_width, const int ch_src)

Get the required buffer size for S8 average pooling function for processors with DSP extension.

Unlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. arm_avgpool_s8() depends on that sentinel being non-zero, since it reads a non-zero size as “ctx->buf is required” before touching the accumulator buffer.

Parameters of arm_avgpool_s8_get_buffer_size_dsp
NameTypeDirectionDescription
dim_dst_widthconst intinoutput tensor dimension
ch_srcconst intinnumber of input tensor channels
Returns of arm_avgpool_s8_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t
function

Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:5684

int32_t arm_avgpool_s8_get_buffer_size_mve(const int dim_dst_width, const int ch_src)

Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case.

Unlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. arm_avgpool_s8() depends on that sentinel being non-zero, since it reads a non-zero size as “ctx->buf is required” before touching the accumulator buffer.

Parameters of arm_avgpool_s8_get_buffer_size_mve
NameTypeDirectionDescription
dim_dst_widthconst intinoutput tensor dimension
ch_srcconst intinnumber of input tensor channels
Returns of arm_avgpool_s8_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t
function

s16 average pooling function.

Include/arm_nnfunctions.h:5721

arm_cmsis_nn_status arm_avgpool_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params *pool_params,
const cmsis_nn_dims *input_dims,
const int16_t *input_data,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int16_t *output_data
)

s16 average pooling function.

  • Supported Framework: TensorFlow Lite
Parameters of arm_avgpool_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outFunction context. Size ctx->buf with arm_avgpool_s16_get_buffer_size(output_dims->w, input_dims->c). Note that it takes the output width and the input channel count rather than a `cmsis_nn_dims`. `arm_avgpool_s16_get_buffer_size_dsp()` and `arm_avgpool_s16_get_buffer_size_mve()` size the same buffer for a specific target. It returns 0 where no buffer is needed, in which case ctx->buf may be NULL. The caller is expected to clear the buffer, if applicable, for security reasons.
pool_paramsconst cmsis_nn_pool_params *inPooling parameters
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN]
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [H, W] Argument N and C are not used.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN.
output_dataint16_t *outOutput data pointer. Data type: int16
Returns of arm_avgpool_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns (an extent of 0 or less), which writes nothing and does not use ctx `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments, including a negative channel count, a pooling window that does not overlap the input, window positions (output index times stride minus padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t, or, on builds that use the buffer, a NULL ctx, or a NULL ctx->buf where the sizer asks for a buffer. Nothing is written to output_data then.
function

Get the required buffer size for S16 average pooling function.

Include/arm_nnfunctions.h:5741

int32_t arm_avgpool_s16_get_buffer_size(const int dim_dst_width, const int ch_src)

Get the required buffer size for S16 average pooling function.

As in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.

Parameters of arm_avgpool_s16_get_buffer_size
NameTypeDirectionDescription
dim_dst_widthconst intinoutput tensor dimension
ch_srcconst intinnumber of input tensor channels
Returns of arm_avgpool_s16_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t
function

Get the required buffer size for S16 average pooling function for processors with DSP extension.

Include/arm_nnfunctions.h:5753

int32_t arm_avgpool_s16_get_buffer_size_dsp(const int dim_dst_width, const int ch_src)

Get the required buffer size for S16 average pooling function for processors with DSP extension.

As in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.

Parameters of arm_avgpool_s16_get_buffer_size_dsp
NameTypeDirectionDescription
dim_dst_widthconst intinoutput tensor dimension
ch_srcconst intinnumber of input tensor channels
Returns of arm_avgpool_s16_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t
function

Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:5766

int32_t arm_avgpool_s16_get_buffer_size_mve(const int dim_dst_width, const int ch_src)

Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case.

As in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.

Parameters of arm_avgpool_s16_get_buffer_size_mve
NameTypeDirectionDescription
dim_dst_widthconst intinoutput tensor dimension
ch_srcconst intinnumber of input tensor channels
Returns of arm_avgpool_s16_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t
function

s8 max pooling function.

Include/arm_nnfunctions.h:5798

arm_cmsis_nn_status arm_max_pool_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params *pool_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 max pooling function.

  • Supported Framework: TensorFlow Lite
Parameters of arm_max_pool_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_max_pool_s8_get_buffer_size(). Max pooling needs no accumulator scratch, unlike `arm_avgpool_s8()`, whose sizer does not describe this argument.
pool_paramsconst cmsis_nn_pool_params *inPooling parameters
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN]
input_dataconst int8_t *inInput (activation) data pointer. The input tensor must not overlap with the output tensor. Data type: int8
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [H, W] Argument N and C are not used.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN.
output_dataint8_t *outOutput data pointer. Data type: int8
Returns of arm_max_pool_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns, which writes nothing `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments: a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to output_data then.
function

s16 max pooling function.

Include/arm_nnfunctions.h:5836

arm_cmsis_nn_status arm_max_pool_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_pool_params *pool_params,
const cmsis_nn_dims *input_dims,
const int16_t *src,
const cmsis_nn_dims *filter_dims,
const cmsis_nn_dims *output_dims,
int16_t *dst
)

s16 max pooling function.

  • Supported Framework: TensorFlow Lite
Parameters of arm_max_pool_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inFunction context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_max_pool_s16_get_buffer_size(). Max pooling needs no accumulator scratch, unlike `arm_avgpool_s16()`, whose sizer does not describe this argument.
pool_paramsconst cmsis_nn_pool_params *inPooling parameters
input_dimsconst cmsis_nn_dims *inInput (activation) tensor dimensions. Format: [H, W, C_IN]
srcconst int16_t *inInput (activation) data pointer. The input tensor must not overlap with the output tensor. Data type: int16
filter_dimsconst cmsis_nn_dims *inFilter tensor dimensions. Format: [H, W] Argument N and C are not used.
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN.
dstint16_t *in, outOutput data pointer. Data type: int16
Returns of arm_max_pool_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns, which writes nothing `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments: a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then.
function

S8 softmax function.

Include/arm_nnfunctions.h:5864

void arm_softmax_s8(
const int8_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const int32_t diff_min,
int8_t *output
)

S8 softmax function.

Parameters of arm_softmax_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input tensor
num_rowsconst int32_tinNumber of rows in the input tensor
row_sizeconst int32_tinNumber of elements in each input row
multconst int32_tinInput quantization multiplier
shiftconst int32_tinInput quantization shift within the range [0, 31]
diff_minconst int32_tinMinimum difference with max in row. Used to check if the quantized exponential operation can be performed
outputint8_t *outPointer to the output tensor
function

S8 to s16 softmax function.

Include/arm_nnfunctions.h:5886

void arm_softmax_s8_s16(
const int8_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const int32_t diff_min,
int16_t *output
)

S8 to s16 softmax function.

Parameters of arm_softmax_s8_s16
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input tensor
num_rowsconst int32_tinNumber of rows in the input tensor
row_sizeconst int32_tinNumber of elements in each input row
multconst int32_tinInput quantization multiplier
shiftconst int32_tinInput quantization shift within the range [0, 31]
diff_minconst int32_tinMinimum difference with max in row. Used to check if the quantized exponential operation can be performed
outputint16_t *outPointer to the output tensor
function

S16 softmax function.

Include/arm_nnfunctions.h:5915

arm_cmsis_nn_status arm_softmax_s16(
const int16_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const cmsis_nn_softmax_lut_s16 *softmax_params,
int16_t *output
)

S16 softmax function.

Parameters of arm_softmax_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input tensor
num_rowsconst int32_tinNumber of rows in the input tensor
row_sizeconst int32_tinNumber of elements in each input row
multconst int32_tinInput quantization multiplier
shiftconst int32_tinInput quantization shift within the range [0, 31]
softmax_paramsconst cmsis_nn_softmax_lut_s16 *inSoftmax s16 layer parameters with two pointers to LUTs speficied below. For indexing the high 9 bits are used and 7 remaining for interpolation. That means 512 entries for the 9-bit indexing and 1 extra for interpolation, i.e. 513 values for each LUT. - Lookup table for exp(x), where x uniform distributed between [-10.0 , 0.0] - Lookup table for 1 / (1 + x), where x uniform distributed between [0.0 , 1.0]
outputint16_t *outPointer to the output tensor
Returns of arm_softmax_s16
Description
The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation
function

U8 softmax function.

Include/arm_nnfunctions.h:5938

void arm_softmax_u8(
const uint8_t *input,
const int32_t num_rows,
const int32_t row_size,
const int32_t mult,
const int32_t shift,
const int32_t diff_min,
uint8_t *output
)

U8 softmax function.

Parameters of arm_softmax_u8
NameTypeDirectionDescription
inputconst uint8_t *inPointer to the input tensor
num_rowsconst int32_tinNumber of rows in the input tensor
row_sizeconst int32_tinNumber of elements in each input row
multconst int32_tinInput quantization multiplier
shiftconst int32_tinInput quantization shift within the range [0, 31]
diff_minconst int32_tinMinimum difference with max in row. Used to check if the quantized exponential operation can be performed
outputuint8_t *outPointer to the output tensor
function

Reshape a s8 vector into another with different shape.

Include/arm_nnfunctions.h:5960

void arm_reshape_s8(const int8_t *input, int8_t *output, const uint32_t total_size)

Reshape a s8 vector into another with different shape.

Parameters of arm_reshape_s8
NameTypeDirectionDescription
inputconst int8_t *inpoints to the s8 input vector
outputint8_t *outpoints to the s8 output vector
total_sizeconst uint32_tintotal size of the input and output vectors in bytes
function

Nearest neighbor resize function for s8 data.

Include/arm_nnfunctions.h:5982

arm_cmsis_nn_status arm_resize_nearest_neighbor_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_resize_params *resize_params,
const cmsis_nn_dims *input_shape,
const int8_t *input_data,
const cmsis_nn_dims *output_size_shape,
const int32_t *output_size_data,
const cmsis_nn_dims *output_shape,
int8_t *output_data
)

Nearest neighbor resize function for s8 data.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_resize_nearest_neighbor_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inPointer to the context buffer. The buffer must hold at least (output_height + output_width) int32_t elements.
resize_paramsconst cmsis_nn_resize_params *inResize parameters
input_shapeconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C]
input_dataconst int8_t *inPointer to input tensor data
output_size_shapeconst cmsis_nn_dims *inOutput size tensor dimensions
output_size_dataconst int32_t *inOutput size tensor data
output_shapeconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C]
output_dataint8_t *outPointer to output tensor data
Returns of arm_resize_nearest_neighbor_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Nearest neighbor resize function for s16 data.

Include/arm_nnfunctions.h:6011

arm_cmsis_nn_status arm_resize_nearest_neighbor_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_resize_params *resize_params,
const cmsis_nn_dims *input_shape,
const int16_t *input_data,
const cmsis_nn_dims *output_size_shape,
const int32_t *output_size_data,
const cmsis_nn_dims *output_shape,
int16_t *output_data
)

Nearest neighbor resize function for s16 data.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_resize_nearest_neighbor_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inPointer to the context buffer. The buffer must hold at least (output_height + output_width) int32_t elements.
resize_paramsconst cmsis_nn_resize_params *inResize parameters
input_shapeconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C]
input_dataconst int16_t *inPointer to input tensor data
output_size_shapeconst cmsis_nn_dims *inOutput size tensor dimensions
output_size_dataconst int32_t *inOutput size tensor data
output_shapeconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H, W, C]
output_dataint16_t *outPointer to output tensor data
Returns of arm_resize_nearest_neighbor_s16
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Space to Depth function for s8 data type.

Include/arm_nnfunctions.h:6036

arm_cmsis_nn_status arm_space_to_depth_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t block_size,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)

Space to Depth function for s8 data type.

  • Supported Framework: TensorFlow Lite
Parameters of arm_space_to_depth_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to the input tensor. Data type: int8
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
block_sizeconst int32_tinBlock size for space to depth transformation
output_dataint8_t *outPointer to the output tensor. Data type: int8
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H/block_size, W/block_size, C_IN*block_size*block_size]
Returns of arm_space_to_depth_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Space to Depth function for s16 data type.

Include/arm_nnfunctions.h:6058

arm_cmsis_nn_status arm_space_to_depth_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t block_size,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)

Space to Depth function for s16 data type.

  • Supported Framework: TensorFlow Lite
Parameters of arm_space_to_depth_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to the input tensor. Data type: int16
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
block_sizeconst int32_tinBlock size for space to depth transformation
output_dataint16_t *outPointer to the output tensor. Data type: int16
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H/block_size, W/block_size, C_IN*block_size*block_size]
Returns of arm_space_to_depth_s16
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Depth to Space function for s8 data type.

Include/arm_nnfunctions.h:6080

arm_cmsis_nn_status arm_depth_to_space_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t block_size,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)

Depth to Space function for s8 data type.

  • Supported Framework: TensorFlow Lite
Parameters of arm_depth_to_space_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to the input tensor. Data type: int8
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
block_sizeconst int32_tinBlock size for depth to space transformation
output_dataint8_t *outPointer to the output tensor. Data type: int8
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H*block_size, W*block_size, C_IN/(block_size*block_size)]
Returns of arm_depth_to_space_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Depth to Space function for s16 data type.

Include/arm_nnfunctions.h:6102

arm_cmsis_nn_status arm_depth_to_space_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t block_size,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)

Depth to Space function for s16 data type.

  • Supported Framework: TensorFlow Lite
Parameters of arm_depth_to_space_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to the input tensor. Data type: int16
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
block_sizeconst int32_tinBlock size for depth to space transformation
output_dataint16_t *outPointer to the output tensor. Data type: int16
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N, H*block_size, W*block_size, C_IN/(block_size*block_size)]
Returns of arm_depth_to_space_s16
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Space to Batch ND function for s8 data type.

Include/arm_nnfunctions.h:6126

arm_cmsis_nn_status arm_space_to_batch_nd_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_tile *block_shape,
const cmsis_nn_dims *pad,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t output_offset
)

Space to Batch ND function for s8 data type.

  • Supported Framework: TensorFlow Lite
Parameters of arm_space_to_batch_nd_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to the input tensor. Data type: int8
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
block_shapeconst cmsis_nn_tile *inBlock shape for space to batch transformation
padconst cmsis_nn_dims *inPadding for height and width. Format: [n->top, h->left, w->bottom, c->right]
output_dataint8_t *outPointer to the output tensor. Data type: int8
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N*block_shape[0]*block_shape[1], (H + pad_top + pad_bottom)/block_shape[0], (W + pad_left + pad_right)/block_shape[1], C_IN]
output_offsetconst int32_tinZero offset for the output tensor
Returns of arm_space_to_batch_nd_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Space to Batch ND function for s16 data type.

Include/arm_nnfunctions.h:6152

arm_cmsis_nn_status arm_space_to_batch_nd_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_tile *block_shape,
const cmsis_nn_dims *pad,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t output_offset
)

Space to Batch ND function for s16 data type.

  • Supported Framework: TensorFlow Lite
Parameters of arm_space_to_batch_nd_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to the input tensor. Data type: int16
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
block_shapeconst cmsis_nn_tile *inBlock shape for space to batch transformation
padconst cmsis_nn_dims *inPadding for height and width. Format: [n->top, h->left, w->bottom, c->right]
output_dataint16_t *outPointer to the output tensor. Data type: int16
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N*block_shape[0]*block_shape[1], (H + pad_top + pad_bottom)/block_shape[0], (W + pad_left + pad_right)/block_shape[1], C_IN]
output_offsetconst int32_tinZero offset for the output tensor. NOT USED. Assume symmetric quantization for s16.
Returns of arm_space_to_batch_nd_s16
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Batch to Space ND function for s8 data type.

Include/arm_nnfunctions.h:6177

arm_cmsis_nn_status arm_batch_to_space_nd_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_tile *block_shape,
const cmsis_nn_dims *crop,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)

Batch to Space ND function for s8 data type.

  • Supported Framework: TensorFlow Lite
Parameters of arm_batch_to_space_nd_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to the input tensor. Data type: int8
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
block_shapeconst cmsis_nn_tile *inBlock shape for batch to space transformation
cropconst cmsis_nn_dims *inCropping for height and width. Format: [n->top, h->left, w->bottom, c->right]
output_dataint8_t *outPointer to the output tensor. Data type: int8
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N/(block_shape[0]*block_shape[1]), H*block_shape[0] - crop_top - crop_bottom, W*block_shape[1] - crop_left - crop_right, C_IN]
Returns of arm_batch_to_space_nd_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Batch to Space ND function for s16 data type.

Include/arm_nnfunctions.h:6201

arm_cmsis_nn_status arm_batch_to_space_nd_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_tile *block_shape,
const cmsis_nn_dims *crop,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)

Batch to Space ND function for s16 data type.

  • Supported Framework: TensorFlow Lite
Parameters of arm_batch_to_space_nd_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to the input tensor. Data type: int16
input_dimsconst cmsis_nn_dims *inInput tensor dimensions. Format: [N, H, W, C_IN]
block_shapeconst cmsis_nn_tile *inBlock shape for batch to space transformation
cropconst cmsis_nn_dims *inCropping for height and width. Format: [n->top, h->left, w->bottom, c->right]
output_dataint16_t *outPointer to the output tensor. Data type: int16
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions. Format: [N/(block_shape[0]*block_shape[1]), H*block_shape[0] - crop_top - crop_bottom, W*block_shape[1] - crop_left - crop_right, C_IN]
Returns of arm_batch_to_space_nd_s16
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Basic transpose function.

Include/arm_nnfunctions.h:6235

arm_cmsis_nn_status arm_transpose_s8(
const int8_t *input_data,
int8_t *const output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const output_dims,
const cmsis_nn_transpose_params *const transpose_params
)

Basic transpose function.

Parameters of arm_transpose_s8
NameTypeDirectionDescription
input_dataconst int8_t *inInput (activation) data pointer. Data type: int8
output_dataint8_t *constoutOutput data pointer. Data type: int8
input_dimsconst cmsis_nn_dims *constinInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
output_dimsconst cmsis_nn_dims *constinOutput tensor dimensions. Format may be arbitrary relative to input format. The output dimension will depend on the permutation dimensions. In other words the out dimensions are the result of applying the permutation to the input dimensions. The first transpose_params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[permutations[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not.
transpose_paramsconst cmsis_nn_transpose_params *constinTranspose parameters. Contains permutation dimensions. num_dims must be in [1, 4] and permutations must be a bijection over [0, num_dims - 1].
Returns of arm_transpose_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

Basic s16 transpose function.

Include/arm_nnfunctions.h:6263

arm_cmsis_nn_status arm_transpose_s16(
const int16_t *input_data,
int16_t *const output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const output_dims,
const cmsis_nn_transpose_params *const transpose_params
)

Basic s16 transpose function.

Parameters of arm_transpose_s16
NameTypeDirectionDescription
input_dataconst int16_t *inInput (activation) data pointer. Data type: int16
output_dataint16_t *constoutOutput data pointer. Data type: int16
input_dimsconst cmsis_nn_dims *constinInput (activation) tensor dimensions. Format: [N, H, W, C_IN]
output_dimsconst cmsis_nn_dims *constinOutput tensor dimensions. Format may be arbitrary relative to input format. The output dimension will depend on the permutation dimensions. In other words the out dimensions are the result of applying the permutation to the input dimensions. The first transpose_params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[permutations[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not.
transpose_paramsconst cmsis_nn_transpose_params *constinTranspose parameters. Contains permutation dimensions. num_dims must be in [1, 4] and permutations must be a bijection over [0, num_dims - 1].
Returns of arm_transpose_s16
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate.

Include/arm_nnfunctions.h:6312

void arm_concatenation_s8_x(
const int8_t *input,
const uint16_t input_x,
const uint16_t input_y,
const uint16_t input_z,
const uint16_t input_w,
int8_t *output,
const uint16_t output_x,
const uint32_t offset_x
)

int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate. The argument offset_x will be used to store the input tensor in the correct position in the output tensor

i.e. offset_x = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_x(&input[i], …, &output, …, …, offset_x) offset_x += input_x[i] }

This function assumes that the output tensor has:

  1. The same height of the input tensor
  2. The same number of channels of the input tensor
  3. The same batch size of the input tensor

Unless specified otherwise, arguments are mandatory.

Input constraints offset_x is less than output_x

Parameters of arm_concatenation_s8_x
NameTypeDirectionDescription
inputconst int8_t *inPointer to input tensor. Input tensor must not overlap with the output tensor.
input_xconst uint16_tinWidth of input tensor
input_yconst uint16_tinHeight of input tensor
input_zconst uint16_tinChannels in input tensor
input_wconst uint16_tinBatch size in input tensor
outputint8_t *outPointer to output tensor. Expected to be at least (input_x * input_y * input_z * input_w) + offset_x bytes.
output_xconst uint16_tinWidth of output tensor
offset_xconst uint32_tinThe offset (in number of elements) on the X axis to start concatenating the input tensor It is user responsibility to provide the correct value
function

int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate.

Include/arm_nnfunctions.h:6359

void arm_concatenation_s8_y(
const int8_t *input,
const uint16_t input_x,
const uint16_t input_y,
const uint16_t input_z,
const uint16_t input_w,
int8_t *output,
const uint16_t output_y,
const uint32_t offset_y
)

int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate. The argument offset_y will be used to store the input tensor in the correct position in the output tensor

i.e. offset_y = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_y(&input[i], …, &output, …, …, offset_y) offset_y += input_y[i] }

This function assumes that the output tensor has:

  1. The same width of the input tensor
  2. The same number of channels of the input tensor
  3. The same batch size of the input tensor

Unless specified otherwise, arguments are mandatory.

Input constraints offset_y is less than output_y

Parameters of arm_concatenation_s8_y
NameTypeDirectionDescription
inputconst int8_t *inPointer to input tensor. Input tensor must not overlap with the output tensor.
input_xconst uint16_tinWidth of input tensor
input_yconst uint16_tinHeight of input tensor
input_zconst uint16_tinChannels in input tensor
input_wconst uint16_tinBatch size in input tensor
outputint8_t *outPointer to output tensor. Expected to be at least (input_z * input_w * input_x * input_y) + offset_y bytes.
output_yconst uint16_tinHeight of output tensor
offset_yconst uint32_tinThe offset on the Y axis to start concatenating the input tensor It is user responsibility to provide the correct value
function

int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate.

Include/arm_nnfunctions.h:6406

void arm_concatenation_s8_z(
const int8_t *input,
const uint16_t input_x,
const uint16_t input_y,
const uint16_t input_z,
const uint16_t input_w,
int8_t *output,
const uint16_t output_z,
const uint32_t offset_z
)

int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate. The argument offset_z will be used to store the input tensor in the correct position in the output tensor

i.e. offset_z = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_z(&input[i], …, &output, …, …, offset_z) offset_z += input_z[i] }

This function assumes that the output tensor has:

  1. The same width of the input tensor
  2. The same height of the input tensor
  3. The same batch size of the input tensor

Unless specified otherwise, arguments are mandatory.

Input constraints offset_z is less than output_z

Parameters of arm_concatenation_s8_z
NameTypeDirectionDescription
inputconst int8_t *inPointer to input tensor. Input tensor must not overlap with output tensor.
input_xconst uint16_tinWidth of input tensor
input_yconst uint16_tinHeight of input tensor
input_zconst uint16_tinChannels in input tensor
input_wconst uint16_tinBatch size in input tensor
outputint8_t *outPointer to output tensor. Expected to be at least (input_x * input_y * input_z * input_w) + offset_z bytes.
output_zconst uint16_tinChannels in output tensor
offset_zconst uint32_tinThe offset on the Z axis to start concatenating the input tensor It is user responsibility to provide the correct value
function

int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to…

Include/arm_nnfunctions.h:6449

void arm_concatenation_s8_w(
const int8_t *input,
const uint16_t input_x,
const uint16_t input_y,
const uint16_t input_z,
const uint16_t input_w,
int8_t *output,
const uint32_t offset_w
)

int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to concatenate. The argument offset_w will be used to store the input tensor in the correct position in the output tensor

i.e. offset_w = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_w(&input[i], …, &output, …, …, offset_w) offset_w += input_w[i] }

This function assumes that the output tensor has:

  1. The same width of the input tensor
  2. The same height of the input tensor
  3. The same number o channels of the input tensor

Unless specified otherwise, arguments are mandatory.

Parameters of arm_concatenation_s8_w
NameTypeDirectionDescription
inputconst int8_t *inPointer to input tensor
input_xconst uint16_tinWidth of input tensor
input_yconst uint16_tinHeight of input tensor
input_zconst uint16_tinChannels in input tensor
input_wconst uint16_tinBatch size in input tensor
outputint8_t *outPointer to output tensor. Expected to be at least input_x * input_y * input_z * input_w bytes.
offset_wconst uint32_tinThe offset on the W axis to start concatenating the input tensor It is user responsibility to provide the correct value
function

int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis

Include/arm_nnfunctions.h:6474

arm_cmsis_nn_status arm_concatenation_s8(
const int8_t *const *input_data,
const int32_t inputs_count,
const int32_t *input_concat_dims,
const int32_t axis,
int8_t *output_data,
const int32_t output_dims,
const int32_t *output_shape
)

int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis

Parameters of arm_concatenation_s8
NameTypeDirectionDescription
input_dataconst int8_t *const *inPointer to input tensors
inputs_countconst int32_tinNumber of input tensors
input_concat_dimsconst int32_t *inDimensions of the input tensors along the target axis
axisconst int32_tinTarget axis to concatenate the input tensors
output_dataint8_t *outPointer to output tensor
output_dimsconst int32_tinOutput tensor dimensions
output_shapeconst int32_t *inOutput tensor shape
Returns of arm_concatenation_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis

Include/arm_nnfunctions.h:6499

arm_cmsis_nn_status arm_concatenation_s16(
const int16_t *const *input_data,
const int32_t inputs_count,
const int32_t *input_concat_dims,
const int32_t axis,
int16_t *output_data,
const int32_t output_dims,
const int32_t *output_shape
)

int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis

Parameters of arm_concatenation_s16
NameTypeDirectionDescription
input_dataconst int16_t *const *inPointer to input tensors
inputs_countconst int32_tinNumber of input tensors
input_concat_dimsconst int32_t *inDimensions of the input tensors along the target axis
axisconst int32_tinTarget axis to concatenate the input tensors
output_dataint16_t *outPointer to output tensor
output_dimsconst int32_tinOutput tensor dimensions
output_shapeconst int32_t *inOutput tensor shape
Returns of arm_concatenation_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis

Include/arm_nnfunctions.h:6524

arm_cmsis_nn_status arm_concatenation_s32(
const int32_t *const *input_data,
const int32_t inputs_count,
const int32_t *input_concat_dims,
const int32_t axis,
int32_t *output_data,
const int32_t output_dims,
const int32_t *output_shape
)

int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis

Parameters of arm_concatenation_s32
NameTypeDirectionDescription
input_dataconst int32_t *const *inPointer to input tensors
inputs_countconst int32_tinNumber of input tensors
input_concat_dimsconst int32_t *inDimensions of the input tensors along the target axis
axisconst int32_tinTarget axis to concatenate the input tensors
output_dataint32_t *outPointer to output tensor
output_dimsconst int32_tinOutput tensor dimensions
output_shapeconst int32_t *inOutput tensor shape
Returns of arm_concatenation_s32
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis

Include/arm_nnfunctions.h:6547

arm_cmsis_nn_status arm_split_s8(
const int8_t *input_data,
const int32_t input_dims,
const int32_t *input_shape,
const int32_t axis,
const int32_t num_splits,
const int32_t *split_dims,
int8_t *const *output_data
)

int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis

Parameters of arm_split_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to the flattened input tensor data.
input_dimsconst int32_tinNumber of dimensions in input_shape.
input_shapeconst int32_t *inArray of length input_dims describing the shape of input_data.
axisconst int32_tinAxis along which to split (0 <= axis < input_dims).
num_splitsconst int32_tinNumber of output tensors to produce.
split_dimsconst int32_t *inArray of length num_splits giving size of each slice along axis.
output_dataint8_t *const *outArray of pointers; output_data[i] points to storage for the i-th output tensor.
Returns of arm_split_s8
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR if split_dims sum mismatch.
function

int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis

Include/arm_nnfunctions.h:6570

arm_cmsis_nn_status arm_split_s16(
const int16_t *input_data,
const int32_t input_dims,
const int32_t *input_shape,
const int32_t axis,
const int32_t num_splits,
const int32_t *split_dims,
int16_t *const *output_data
)

int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis

Parameters of arm_split_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to the flattened input tensor data.
input_dimsconst int32_tinNumber of dimensions in input_shape.
input_shapeconst int32_t *inArray of length input_dims describing the shape of input_data.
axisconst int32_tinAxis along which to split (0 <= axis < input_dims).
num_splitsconst int32_tinNumber of output tensors to produce.
split_dimsconst int32_t *inArray of length num_splits giving size of each slice along axis.
output_dataint16_t *const *outArray of pointers; output_data[i] points to storage for the i-th output tensor.
Returns of arm_split_s16
Description
ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR if split_dims sum mismatch.
function

s8 SVDF function with 8 bit state tensor and 8 bit time weights

Include/arm_nnfunctions.h:6662

arm_cmsis_nn_status arm_svdf_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_context *input_ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_svdf_params *svdf_params,
const cmsis_nn_per_tensor_quant_params *input_quant_params,
const cmsis_nn_per_tensor_quant_params *output_quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *state_dims,
int8_t *state_data,
const cmsis_nn_dims *weights_feature_dims,
const int8_t *weights_feature_data,
const cmsis_nn_dims *weights_time_dims,
const int8_t *weights_time_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 SVDF function with 8 bit state tensor and 8 bit time weights

  1. Supported framework: TensorFlow Lite micro
Parameters of arm_svdf_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inPrecomputed per-feature-batch kernel sums, supplied by the caller. This is an input the function only reads, not scratch it fills: an allocated but unfilled buffer yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. Mandatory on builds with the MVE extension (ARM_MATH_MVEI), where a NULL ctx->buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. Unused on every other build, where ctx->buf may be NULL. Sized by arm_svdf_s8_get_buffer_size(weights_feature_dims): weights_feature_dims->n * sizeof(int32_t) where the sums are used, 0 otherwise. Note this is weights_feature_dims->n, not a filter_dims->c - do not size this buffer with `arm_fully_connected_s8_get_buffer_size()`, which reads a different field and under-allocates. Fill it with arm_vector_sum_s8(ctx->buf, input_dims->h, weights_feature_dims->n, weights_feature_data, -svdf_params->input_offset, 0, NULL) so that entry j holds -input_offset * sum(weights_feature row j). The contents depend only on weights_feature_data and svdf_params->input_offset, so they may be computed once at load time and reused across calls until one of those changes. The buffer is specific to one layer's weights and cannot be shared between layers. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call.
input_ctxconst cmsis_nn_context *inScratch buffer written by this function, holding one int32_t accumulator per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_s8_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(int32_t) bytes, the same figure on every build target. This function does not read input_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons.
output_ctxconst cmsis_nn_context *inScratch buffer written by this function, holding one int32_t accumulator per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_s8_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t) bytes, truncating division, the same figure on every build target. This function does not read output_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons.
svdf_paramsconst cmsis_nn_svdf_params *inSVDF Parameters Range of svdf_params->input_offset : [-128, 127] Range of svdf_params->output_offset : [-128, 127]
input_quant_paramsconst cmsis_nn_per_tensor_quant_params *inInput quantization parameters
output_quant_paramsconst cmsis_nn_per_tensor_quant_params *inOutput quantization parameters
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
input_dataconst int8_t *inPointer to input tensor
state_dimsconst cmsis_nn_dims *inState tensor dimensions
state_dataint8_t *in, outPointer to state tensor
weights_feature_dimsconst cmsis_nn_dims *inWeights (feature) tensor dimensions
weights_feature_dataconst int8_t *inPointer to the weights (feature) tensor
weights_time_dimsconst cmsis_nn_dims *inWeights (time) tensor dimensions
weights_time_dataconst int8_t *inPointer to the weights (time) tensor
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions
bias_dataconst int32_t *inPointer to bias tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
output_dataint8_t *outPointer to the output tensor
Returns of arm_svdf_s8
Description
The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion.
function

s8 SVDF function with 16 bit state tensor and 16 bit time weights

Include/arm_nnfunctions.h:6734

arm_cmsis_nn_status arm_svdf_state_s16_s8(
const cmsis_nn_context *input_ctx,
const cmsis_nn_context *output_ctx,
const cmsis_nn_svdf_params *svdf_params,
const cmsis_nn_per_tensor_quant_params *input_quant_params,
const cmsis_nn_per_tensor_quant_params *output_quant_params,
const cmsis_nn_dims *input_dims,
const int8_t *input_data,
const cmsis_nn_dims *state_dims,
int16_t *state_data,
const cmsis_nn_dims *weights_feature_dims,
const int8_t *weights_feature_data,
const cmsis_nn_dims *weights_time_dims,
const int16_t *weights_time_data,
const cmsis_nn_dims *bias_dims,
const int32_t *bias_data,
const cmsis_nn_dims *output_dims,
int8_t *output_data
)

s8 SVDF function with 16 bit state tensor and 16 bit time weights

  1. Supported framework: TensorFlow Lite micro
Parameters of arm_svdf_state_s16_s8
NameTypeDirectionDescription
input_ctxconst cmsis_nn_context *inScratch buffer written by this function, holding one int32_t accumulator per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_state_s16_s8_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(int32_t) bytes, the same figure on every build target. Note the accumulators are int32_t even though the state tensor is int16_t - this buffer does not shrink with the state width. This function does not read input_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons.
output_ctxconst cmsis_nn_context *inScratch buffer written by this function, holding one int32_t accumulator per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_state_s16_s8_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t) bytes, truncating division, the same figure on every build target. This function does not read output_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons.
svdf_paramsconst cmsis_nn_svdf_params *inSVDF Parameters Range of svdf_params->input_offset : [-128, 127] Range of svdf_params->output_offset : [-128, 127]
input_quant_paramsconst cmsis_nn_per_tensor_quant_params *inInput quantization parameters
output_quant_paramsconst cmsis_nn_per_tensor_quant_params *inOutput quantization parameters
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
input_dataconst int8_t *inPointer to input tensor
state_dimsconst cmsis_nn_dims *inState tensor dimensions
state_dataint16_t *in, outPointer to state tensor
weights_feature_dimsconst cmsis_nn_dims *inWeights (feature) tensor dimensions
weights_feature_dataconst int8_t *inPointer to the weights (feature) tensor
weights_time_dimsconst cmsis_nn_dims *inWeights (time) tensor dimensions
weights_time_dataconst int16_t *inPointer to the weights (time) tensor
bias_dimsconst cmsis_nn_dims *inBias tensor dimensions
bias_dataconst int32_t *inPointer to bias tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
output_dataint8_t *outPointer to the output tensor
Returns of arm_svdf_state_s16_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Get size of the kernel-sum buffer required by armsvdfs8().

Include/arm_nnfunctions.h:6767

int32_t arm_svdf_s8_get_buffer_size(const cmsis_nn_dims *weights_feature_dims)

Get size of the kernel-sum buffer required by arm_svdf_s8().

For a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. arm_svdf_s8() has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other cmsis_nn_dims of that call can size it - in particular arm_fully_connected_s8_get_buffer_size() reads a different field and under-allocates. See arm_svdf_s8() for the buffer’s layout, how to fill it and when it may be reused.

Parameters of arm_svdf_s8_get_buffer_size
NameTypeDirectionDescription
weights_feature_dimsconst cmsis_nn_dims *indimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
Returns of arm_svdf_s8_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t
function

Get size of the kernel-sum buffer required by armsvdfs8() for processors with DSP extension.

Include/arm_nnfunctions.h:6778

int32_t arm_svdf_s8_get_buffer_size_dsp(const cmsis_nn_dims *weights_feature_dims)

Get size of the kernel-sum buffer required by arm_svdf_s8() for processors with DSP extension.

For a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. arm_svdf_s8() has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other cmsis_nn_dims of that call can size it - in particular arm_fully_connected_s8_get_buffer_size() reads a different field and under-allocates. See arm_svdf_s8() for the buffer’s layout, how to fill it and when it may be reused.

Parameters of arm_svdf_s8_get_buffer_size_dsp
NameTypeDirectionDescription
weights_feature_dimsconst cmsis_nn_dims *indimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
Returns of arm_svdf_s8_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t
function

Get size of the kernel-sum buffer required by armsvdfs8() for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:6789

int32_t arm_svdf_s8_get_buffer_size_mve(const cmsis_nn_dims *weights_feature_dims)

Get size of the kernel-sum buffer required by arm_svdf_s8() for Arm(R) Helium Architecture case.

For a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. arm_svdf_s8() has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other cmsis_nn_dims of that call can size it - in particular arm_fully_connected_s8_get_buffer_size() reads a different field and under-allocates. See arm_svdf_s8() for the buffer’s layout, how to fill it and when it may be reused.

Parameters of arm_svdf_s8_get_buffer_size_mve
NameTypeDirectionDescription
weights_feature_dimsconst cmsis_nn_dims *indimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
Returns of arm_svdf_s8_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t
function

Get size of the inputctx staging buffer required by armsvdfs8().

Include/arm_nnfunctions.h:6812

int32_t arm_svdf_s8_input_ctx_get_buffer_size(
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)

Get size of the input_ctx staging buffer required by arm_svdf_s8().

Returns input_dims->n * weights_feature_dims->n * sizeof(int32_t). Unlike arm_svdf_s8_get_buffer_size(), this figure does not vary by build target: arm_svdf_s8() stages this buffer on every build, not only under MVE, so there is no _dsp / _mve pair to choose between and the validation runs on every target.

Parameters of arm_svdf_s8_input_ctx_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
weights_feature_dimsconst cmsis_nn_dims *inWeights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
Returns of arm_svdf_s8_input_ctx_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t
function

Get size of the outputctx staging buffer required by armsvdfs8().

Include/arm_nnfunctions.h:6842

int32_t arm_svdf_s8_output_ctx_get_buffer_size(
const cmsis_nn_svdf_params *svdf_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)

Get size of the output_ctx staging buffer required by arm_svdf_s8().

Returns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t). The division truncates, matching the kernel’s own unit count. As with arm_svdf_s8_input_ctx_get_buffer_size(), the figure is the same on every build target and the validation runs on every target.

Parameters of arm_svdf_s8_output_ctx_get_buffer_size
NameTypeDirectionDescription
svdf_paramsconst cmsis_nn_svdf_params *inSVDF parameters; only svdf_params->rank is read
input_dimsconst cmsis_nn_dims *inInput tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
weights_feature_dimsconst cmsis_nn_dims *inWeights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
Returns of arm_svdf_s8_output_ctx_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any pointer is NULL, if svdf_params->rank is zero, negative or outside int16_t range, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t
function

Get size of the inputctx staging buffer required by armsvdfstates16s8().

Include/arm_nnfunctions.h:6855

int32_t arm_svdf_state_s16_s8_input_ctx_get_buffer_size(
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)

Get size of the input_ctx staging buffer required by arm_svdf_state_s16_s8().

Returns input_dims->n * weights_feature_dims->n * sizeof(int32_t). Unlike arm_svdf_s8_get_buffer_size(), this figure does not vary by build target: arm_svdf_s8() stages this buffer on every build, not only under MVE, so there is no _dsp / _mve pair to choose between and the validation runs on every target.

Returns input_dims->n * weights_feature_dims->n * sizeof(int32_t) - the same figure as arm_svdf_s8_input_ctx_get_buffer_size() for the same shape. The accumulators are int32_t even though arm_svdf_state_s16_s8() carries an int16_t state tensor, so this buffer does not shrink with the state width.

Parameters of arm_svdf_state_s16_s8_input_ctx_get_buffer_size
NameTypeDirectionDescription
input_dimsconst cmsis_nn_dims *inInput tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
weights_feature_dimsconst cmsis_nn_dims *inWeights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
Returns of arm_svdf_state_s16_s8_input_ctx_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t
function

Get size of the outputctx staging buffer required by armsvdfstates16s8().

Include/arm_nnfunctions.h:6865

int32_t arm_svdf_state_s16_s8_output_ctx_get_buffer_size(
const cmsis_nn_svdf_params *svdf_params,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *weights_feature_dims
)

Get size of the output_ctx staging buffer required by arm_svdf_state_s16_s8().

Returns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t). The division truncates, matching the kernel’s own unit count. As with arm_svdf_s8_input_ctx_get_buffer_size(), the figure is the same on every build target and the validation runs on every target.

Returns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t), truncating division - the same figure as arm_svdf_s8_output_ctx_get_buffer_size() for the same shape.

Parameters of arm_svdf_state_s16_s8_output_ctx_get_buffer_size
NameTypeDirectionDescription
svdf_paramsconst cmsis_nn_svdf_params *inSVDF parameters; only svdf_params->rank is read
input_dimsconst cmsis_nn_dims *inInput tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
weights_feature_dimsconst cmsis_nn_dims *inWeights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`
Returns of arm_svdf_state_s16_s8_output_ctx_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if any pointer is NULL, if svdf_params->rank is zero, negative or outside int16_t range, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t
function

LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias.

Include/arm_nnfunctions.h:6892

arm_cmsis_nn_status arm_lstm_unidirectional_s8(
const int8_t *input,
int8_t *output,
const cmsis_nn_lstm_params *params,
cmsis_nn_lstm_context *buffers
)

LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_lstm_unidirectional_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to input data
outputint8_t *outPointer to output data
paramsconst cmsis_nn_lstm_params *inStruct containing all information about the lstm operator, see arm_nn_types.
bufferscmsis_nn_lstm_context *in, outStruct containing pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. Size temp1 with `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` and temp2 with `arm_lstm_unidirectional_s8_temp2_get_buffer_size()` - both hold int16_t gate vectors even though the layer datatype is s8, so sizing them in s8 elements under-allocates by half.
Returns of arm_lstm_unidirectional_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias.

Include/arm_nnfunctions.h:6914

arm_cmsis_nn_status arm_lstm_unidirectional_s16(
const int16_t *input,
int16_t *output,
const cmsis_nn_lstm_params *params,
cmsis_nn_lstm_context *buffers
)

LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_lstm_unidirectional_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to input data
outputint16_t *outPointer to output data
paramsconst cmsis_nn_lstm_params *inStruct containing all information about the lstm operator, see arm_nn_types.
bufferscmsis_nn_lstm_context *in, outStruct containing pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. Size temp1 with `arm_lstm_unidirectional_s16_temp1_get_buffer_size()` and temp2 with `arm_lstm_unidirectional_s16_temp2_get_buffer_size()`.
Returns of arm_lstm_unidirectional_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Get size of the temp1 scratch buffer required by armlstmunidirectionals8().

Include/arm_nnfunctions.h:6940

int32_t arm_lstm_unidirectional_s8_temp1_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)

Get size of the temp1 scratch buffer required by arm_lstm_unidirectional_s8().

Parameters of arm_lstm_unidirectional_s8_temp1_get_buffer_size
NameTypeDirectionDescription
lstm_paramsconst cmsis_nn_lstm_params *inLSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read.
Returns of arm_lstm_unidirectional_s8_temp1_get_buffer_size
Description
Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target.
function

Get size of the temp2 scratch buffer required by armlstmunidirectionals8().

Include/arm_nnfunctions.h:6950

int32_t arm_lstm_unidirectional_s8_temp2_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)

Get size of the temp2 scratch buffer required by arm_lstm_unidirectional_s8().

Parameters of arm_lstm_unidirectional_s8_temp2_get_buffer_size
NameTypeDirectionDescription
lstm_paramsconst cmsis_nn_lstm_params *inLSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read.
Returns of arm_lstm_unidirectional_s8_temp2_get_buffer_size
Description
Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target.
Required buffer size in bytes: the same figure as `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` for the same params. temp2 stages the cell-gate vector and the tanh(cell_state) vector, both of the same extent as the gate vectors staged in temp1.
function

Get size of the temp1 scratch buffer required by armlstmunidirectionals16().

Include/arm_nnfunctions.h:6961

int32_t arm_lstm_unidirectional_s16_temp1_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)

Get size of the temp1 scratch buffer required by arm_lstm_unidirectional_s16().

Parameters of arm_lstm_unidirectional_s16_temp1_get_buffer_size
NameTypeDirectionDescription
lstm_paramsconst cmsis_nn_lstm_params *inLSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read.
Returns of arm_lstm_unidirectional_s16_temp1_get_buffer_size
Description
Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target.
Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t) - the same figure as `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` for the same params, since both layer datatypes stage int16_t gate vectors.
function

Get size of the temp2 scratch buffer required by armlstmunidirectionals16().

Include/arm_nnfunctions.h:6970

int32_t arm_lstm_unidirectional_s16_temp2_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)

Get size of the temp2 scratch buffer required by arm_lstm_unidirectional_s16().

Parameters of arm_lstm_unidirectional_s16_temp2_get_buffer_size
NameTypeDirectionDescription
lstm_paramsconst cmsis_nn_lstm_params *inLSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read.
Returns of arm_lstm_unidirectional_s16_temp2_get_buffer_size
Description
Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target.
Required buffer size in bytes: the same figure as `arm_lstm_unidirectional_s16_temp1_get_buffer_size()` for the same params.
function

Batch matmul function with 8 bit input and output.

Include/arm_nnfunctions.h:7017

arm_cmsis_nn_status arm_batch_matmul_s8(
const cmsis_nn_context *ctx,
const cmsis_nn_bmm_params *bmm_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_lhs_dims,
const int8_t *input_lhs,
const cmsis_nn_dims *input_rhs_dims,
const int8_t *input_rhs,
const cmsis_nn_dims *output_dims,
int8_t *output
)

Batch matmul function with 8 bit input and output.

  1. Supported framework: TensorFlow Lite Micro
  2. Performs row * row matrix multiplication with the RHS transposed.
Parameters of arm_batch_matmul_s8
NameTypeDirectionDescription
ctxconst cmsis_nn_context *in, outTemporary scratch buffer for the per-row kernel sums of the RHS. Mandatory on builds with the MVE extension (ARM_MATH_MVEI), where a NULL ctx->buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. Unused on every other build, where ctx->buf may be NULL. Sized by arm_batch_matmul_s8_get_buffer_size(input_rhs_dims) - pass the same input_rhs_dims given below. That is input_rhs_dims->w * sizeof(int32_t) where the sums are used, 0 otherwise. Do not size this buffer with `arm_fully_connected_s8_get_buffer_size()`: it reads a different field, and an allocation short of input_rhs_dims->w words is written past its end. The function fills the buffer itself before each use, so the caller does not need to initialize it. ctx->buf must be aligned to sizeof(int32_t). If ctx->size is non-zero it is validated against the requirement and a buffer too small is rejected with ARM_CMSIS_NN_ARG_ERROR; a ctx->size of 0 skips that check. A negative input_rhs_dims->w, or one large enough that the required size exceeds INT32_MAX, is rejected with ARM_CMSIS_NN_ARG_ERROR regardless of ctx->size. The caller is expected to clear the buffer, if applicable, for security reasons.
bmm_paramsconst cmsis_nn_bmm_params *inBatch matmul Parameters Adjoint flags are currently unused and do not transpose either input; callers must supply the tensors in the layouts described below.
quant_paramsconst cmsis_nn_per_tensor_quant_params *inQuantization parameters
input_lhs_dimsconst cmsis_nn_dims *inInput lhs tensor dimensions. This s8 function treats w as the row count and c as the inner dimension. This differs from `arm_batch_matmul_f32()`, so its dimension mapping must not be reused here.
input_lhsconst int8_t *inPointer to input tensor
input_rhs_dimsconst cmsis_nn_dims *inInput rhs tensor dimensions. The RHS must already be transposed, with w as its row count and c equal to input_lhs_dims->c.
input_rhsconst int8_t *inPointer to transposed input tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
outputint8_t *outPointer to the output tensor
Returns of arm_batch_matmul_s8
Description
The function returns one of the following: - `ARM_CMSIS_NN_ARG_ERROR` if an MVE build receives an invalid context, a negative or unrepresentable RHS row count, or a declared context size below the requirement. - `ARM_CMSIS_NN_SUCCESS` on success.
function

Batch matmul function with 16 bit input and output.

Include/arm_nnfunctions.h:7057

arm_cmsis_nn_status arm_batch_matmul_s16(
const cmsis_nn_context *ctx,
const cmsis_nn_bmm_params *bmm_params,
const cmsis_nn_per_tensor_quant_params *quant_params,
const cmsis_nn_dims *input_lhs_dims,
const int16_t *input_lhs,
const cmsis_nn_dims *input_rhs_dims,
const int16_t *input_rhs,
const cmsis_nn_dims *output_dims,
int16_t *output
)

Batch matmul function with 16 bit input and output.

  1. Supported framework: TensorFlow Lite Micro
  2. Performs row * row matrix multiplication with the RHS transposed.
Parameters of arm_batch_matmul_s16
NameTypeDirectionDescription
ctxconst cmsis_nn_context *inUnused: this function requires no scratch buffer and does not read or write ctx on any build, so ctx->buf may be NULL. Retained for signature compatibility with `arm_batch_matmul_s8()`. There is deliberately no arm_batch_matmul_s16_get_buffer_size(); in particular `arm_fully_connected_s8_get_buffer_size()` is not the sizer for this argument. If a real buffer is passed, the caller is expected to clear it, if applicable, for security reasons.
bmm_paramsconst cmsis_nn_bmm_params *inBatch matmul Parameters Adjoint flags are currently unused.
quant_paramsconst cmsis_nn_per_tensor_quant_params *inQuantization parameters
input_lhs_dimsconst cmsis_nn_dims *inInput lhs tensor dimensions. This should be NHWC where LHS.C = RHS.C
input_lhsconst int16_t *inPointer to input tensor
input_rhs_dimsconst cmsis_nn_dims *inInput lhs tensor dimensions. This is expected to be transposed so should be NHWC where LHS.C = RHS.C
input_rhsconst int16_t *inPointer to transposed input tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
outputint16_t *outPointer to the output tensor
Returns of arm_batch_matmul_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Get size of the scratch buffer required by armbatchmatmuls8().

Include/arm_nnfunctions.h:7082

int32_t arm_batch_matmul_s8_get_buffer_size(const cmsis_nn_dims *input_rhs_dims)

Get size of the scratch buffer required by arm_batch_matmul_s8().

For a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with arm_fully_connected_s8_get_buffer_size(), which reads .c) writes past the allocation whenever the rhs has more rows than columns. arm_batch_matmul_s16() needs no scratch buffer and so has no corresponding sizer.

Parameters of arm_batch_matmul_s8_get_buffer_size
NameTypeDirectionDescription
input_rhs_dimsconst cmsis_nn_dims *indimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()`
Returns of arm_batch_matmul_s8_get_buffer_size
Description
The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t
function

Get size of the scratch buffer required by armbatchmatmuls8() for processors with DSP extension.

Include/arm_nnfunctions.h:7093

int32_t arm_batch_matmul_s8_get_buffer_size_dsp(const cmsis_nn_dims *input_rhs_dims)

Get size of the scratch buffer required by arm_batch_matmul_s8() for processors with DSP extension.

For a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with arm_fully_connected_s8_get_buffer_size(), which reads .c) writes past the allocation whenever the rhs has more rows than columns. arm_batch_matmul_s16() needs no scratch buffer and so has no corresponding sizer.

Parameters of arm_batch_matmul_s8_get_buffer_size_dsp
NameTypeDirectionDescription
input_rhs_dimsconst cmsis_nn_dims *indimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()`
Returns of arm_batch_matmul_s8_get_buffer_size_dsp
Description
The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t
function

Get size of the scratch buffer required by armbatchmatmuls8() for Arm(R) Helium Architecture case.

Include/arm_nnfunctions.h:7104

int32_t arm_batch_matmul_s8_get_buffer_size_mve(const cmsis_nn_dims *input_rhs_dims)

Get size of the scratch buffer required by arm_batch_matmul_s8() for Arm(R) Helium Architecture case.

For a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with arm_fully_connected_s8_get_buffer_size(), which reads .c) writes past the allocation whenever the rhs has more rows than columns. arm_batch_matmul_s16() needs no scratch buffer and so has no corresponding sizer.

Parameters of arm_batch_matmul_s8_get_buffer_size_mve
NameTypeDirectionDescription
input_rhs_dimsconst cmsis_nn_dims *indimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()`
Returns of arm_batch_matmul_s8_get_buffer_size_mve
Description
The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t
function

Expands the size of the input by adding constant values before and after the data, in all dimensions.

Include/arm_nnfunctions.h:7124

arm_cmsis_nn_status arm_pad_s8(
const int8_t *input,
int8_t *output,
const int8_t pad_value,
const cmsis_nn_dims *input_size,
const cmsis_nn_dims *pre_pad,
const cmsis_nn_dims *post_pad
)

Expands the size of the input by adding constant values before and after the data, in all dimensions.

Parameters of arm_pad_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to input data
outputint8_t *outPointer to output data
pad_valueconst int8_tinValue to pad with
input_sizeconst cmsis_nn_dims *inInput tensor dimensions
pre_padconst cmsis_nn_dims *inPadding to apply before data in each dimension
post_padconst cmsis_nn_dims *inPadding to apply after data in each dimension
Returns of arm_pad_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Expands the size of the input by adding constant values before and after the data, in all dimensions.

Include/arm_nnfunctions.h:7144

arm_cmsis_nn_status arm_pad_s16(
const int16_t *input,
int16_t *output,
const int16_t pad_value,
const cmsis_nn_dims *input_size,
const cmsis_nn_dims *pre_pad,
const cmsis_nn_dims *post_pad
)

Expands the size of the input by adding constant values before and after the data, in all dimensions.

Parameters of arm_pad_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to input data
outputint16_t *outPointer to output data
pad_valueconst int16_tinValue to pad with
input_sizeconst cmsis_nn_dims *inInput tensor dimensions
pre_padconst cmsis_nn_dims *inPadding to apply before data in each dimension
post_padconst cmsis_nn_dims *inPadding to apply after data in each dimension
Returns of arm_pad_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Computes the mean of the input tensor along the specified axis.

Include/arm_nnfunctions.h:7173

arm_cmsis_nn_status arm_mean_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t input_offset,
const cmsis_nn_dims *axis_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift
)

Computes the mean of the input tensor along the specified axis. The output multipler and shift must have the output count folded into them.

Parameters of arm_mean_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
input_offsetconst int32_tinInput offset
axis_dimsconst cmsis_nn_dims *inAxis dimensions to compute mean over
output_dataint8_t *outPointer to output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
out_offsetconst int32_tinOutput offset
out_multconst int32_tinOutput quantization multiplier
out_shiftconst int32_tinOutput quantization shift
Returns of arm_mean_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Computes the mean of the input tensor along the specified axis.

Include/arm_nnfunctions.h:7200

arm_cmsis_nn_status arm_mean_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t input_offset,
const cmsis_nn_dims *axis_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims,
const int32_t out_offset,
const int32_t out_mult,
const int32_t out_shift
)

Computes the mean of the input tensor along the specified axis. The output multipler and shift must have the output count folded into them.

Parameters of arm_mean_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
input_offsetconst int32_tinInput offset
axis_dimsconst cmsis_nn_dims *inAxis dimensions to compute mean over
output_dataint16_t *outPointer to output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
out_offsetconst int32_tinOutput offset
out_multconst int32_tinOutput quantization multiplier
out_shiftconst int32_tinOutput quantization shift
Returns of arm_mean_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Compute ArgMax indices of an s8 tensor along a specific axis.

Include/arm_nnfunctions.h:7221

arm_cmsis_nn_status arm_argmax_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t axis,
int32_t *output_data
)

Compute ArgMax indices of an s8 tensor along a specific axis.

Parameters of arm_argmax_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions (NHWC layout)
axisconst int32_tinReduction axis in range [0, 3]
output_dataint32_t *outPointer to output indices (int32_t)
Returns of arm_argmax_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`.
function

Compute ArgMin indices of an s8 tensor along a specific axis.

Include/arm_nnfunctions.h:7234

arm_cmsis_nn_status arm_argmin_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t axis,
int32_t *output_data
)

Compute ArgMin indices of an s8 tensor along a specific axis.

Parameters of arm_argmin_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions (NHWC layout)
axisconst int32_tinReduction axis in range [0, 3]
output_dataint32_t *outPointer to output indices (int32_t)
Returns of arm_argmin_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`.
function

Compute ArgMax indices of an s16 tensor along a specific axis.

Include/arm_nnfunctions.h:7247

arm_cmsis_nn_status arm_argmax_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t axis,
int32_t *output_data
)

Compute ArgMax indices of an s16 tensor along a specific axis.

Parameters of arm_argmax_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions (NHWC layout)
axisconst int32_tinReduction axis in range [0, 3]
output_dataint32_t *outPointer to output indices (int32_t)
Returns of arm_argmax_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`.
function

Compute ArgMin indices of an s16 tensor along a specific axis.

Include/arm_nnfunctions.h:7260

arm_cmsis_nn_status arm_argmin_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t axis,
int32_t *output_data
)

Compute ArgMin indices of an s16 tensor along a specific axis.

Parameters of arm_argmin_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions (NHWC layout)
axisconst int32_tinReduction axis in range [0, 3]
output_dataint32_t *outPointer to output indices (int32_t)
Returns of arm_argmin_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`.
function

Computes the max of the input tensor along the specified axis.

Include/arm_nnfunctions.h:7274

arm_cmsis_nn_status arm_reduce_max_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)

Computes the max of the input tensor along the specified axis.

Parameters of arm_reduce_max_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
axis_dimsconst cmsis_nn_dims *inAxis dimensions to compute mean over
output_dataint8_t *outPointer to output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_reduce_max_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Computes the max of the input tensor along the specified axis.

Include/arm_nnfunctions.h:7292

arm_cmsis_nn_status arm_reduce_max_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)

Computes the max of the input tensor along the specified axis.

Parameters of arm_reduce_max_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
axis_dimsconst cmsis_nn_dims *inAxis dimensions to compute mean over
output_dataint16_t *outPointer to output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_reduce_max_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Computes the min of the input tensor along the specified axis.

Include/arm_nnfunctions.h:7310

arm_cmsis_nn_status arm_reduce_min_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)

Computes the min of the input tensor along the specified axis.

Parameters of arm_reduce_min_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
axis_dimsconst cmsis_nn_dims *inAxis dimensions to compute mean over
output_dataint8_t *outPointer to output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_reduce_min_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Computes the min of the input tensor along the specified axis.

Include/arm_nnfunctions.h:7328

arm_cmsis_nn_status arm_reduce_min_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const cmsis_nn_dims *axis_dims,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)

Computes the min of the input tensor along the specified axis.

Parameters of arm_reduce_min_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to input tensor
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
axis_dimsconst cmsis_nn_dims *inAxis dimensions to compute mean over
output_dataint16_t *outPointer to output tensor
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_reduce_min_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Quantize a floating-point array into int8t format.

Include/arm_nnfunctions.h:7356

arm_cmsis_nn_status arm_quantize_f32_s8(
const float *input,
int8_t *output,
int32_t size,
int32_t zero_point,
float scale
)

Quantize a floating-point array into int8_t format.

Parameters of arm_quantize_f32_s8
NameTypeDirectionDescription
inputconst float *inPointer to the input float array.
outputint8_t *outPointer to the output int8_t array.
sizeint32_tinNumber of elements in the arrays.
zero_pointint32_tinZero point (offset) to apply during quantization.
scalefloatinScale factor to apply during quantization. Must be a positive finite number. A scale that is zero, negative, NaN, Inf, or small enough that its reciprocal overflows is unsupported, and the result is then unspecified: the scalar and Helium legs are not guaranteed to agree for such a scale. Denormal inputs are likewise unspecified: the Helium leg flushes them to zero, so the two legs are not guaranteed to agree for a denormal input at any scale. Screen denormals if you need them to match.
Returns of arm_quantize_f32_s8
Description
ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when `zero_point` lies outside the int8_t range. Values round half away from zero and saturate to the int8_t range after the zero point is applied; NaN maps to `zero_point`.
function

Quantize a floating-point array into int16t format.

Include/arm_nnfunctions.h:7376

arm_cmsis_nn_status arm_quantize_f32_s16(
const float *input,
int16_t *output,
int32_t size,
int32_t zero_point,
float scale
)

Quantize a floating-point array into int16_t format.

Parameters of arm_quantize_f32_s16
NameTypeDirectionDescription
inputconst float *inPointer to the input float array.
outputint16_t *outPointer to the output int16_t array.
sizeint32_tinNumber of elements in the arrays.
zero_pointint32_tinZero point (offset) to apply during quantization.
scalefloatinScale factor to apply during quantization. Must be a positive finite number. A scale that is zero, negative, NaN, Inf, or small enough that its reciprocal overflows is unsupported, and the result is then unspecified: the scalar and Helium legs are not guaranteed to agree for such a scale. Denormal inputs are likewise unspecified: the Helium leg flushes them to zero, so the two legs are not guaranteed to agree for a denormal input at any scale. Screen denormals if you need them to match.
Returns of arm_quantize_f32_s16
Description
ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when `zero_point` lies outside the int16_t range. Values round half away from zero and saturate to the int16_t range after the zero point is applied; NaN maps to `zero_point`.
function

Requantize an int8t array to another int8t range with a different scale.

Include/arm_nnfunctions.h:7390

arm_cmsis_nn_status arm_requantize_s8_s8(
const int8_t *input,
int8_t *output,
int32_t size,
int32_t effective_scale_multiplier,
int32_t effective_scale_shift,
int32_t input_zeropoint,
int32_t output_zeropoint
)

Requantize an int8_t array to another int8_t range with a different scale.

Parameters of arm_requantize_s8_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input int8_t array.
outputint8_t *outPointer to the output int8_t array.
sizeint32_tinNumber of elements in the arrays.
effective_scale_multiplierint32_tinMultiplier used for the scaling operation.
effective_scale_shiftint32_tinRight or left shift (depending on sign) applied after the multiplier.
input_zeropointint32_tinZero point of the input data.
output_zeropointint32_tinZero point of the output data.
Returns of arm_requantize_s8_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Requantize an int16t array to another int16t range with a different scale.

Include/arm_nnfunctions.h:7410

arm_cmsis_nn_status arm_requantize_s16_s16(
const int16_t *input,
int16_t *output,
int32_t size,
int32_t effective_scale_multiplier,
int32_t effective_scale_shift,
int32_t input_zeropoint,
int32_t output_zeropoint
)

Requantize an int16_t array to another int16_t range with a different scale.

Parameters of arm_requantize_s16_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input int16_t array.
outputint16_t *outPointer to the output int16_t array.
sizeint32_tinNumber of elements in the arrays.
effective_scale_multiplierint32_tinMultiplier used for the scaling operation.
effective_scale_shiftint32_tinRight or left shift (depending on sign) applied after the multiplier.
input_zeropointint32_tinZero point of the input data.
output_zeropointint32_tinZero point of the output data.
Returns of arm_requantize_s16_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Dequantize an int8t array back to floating-point format.

Include/arm_nnfunctions.h:7429

arm_cmsis_nn_status arm_dequantize_s8_f32(
const int8_t *input,
float *output,
int32_t size,
int32_t zero_point,
float scale
)

Dequantize an int8_t array back to floating-point format.

Parameters of arm_dequantize_s8_f32
NameTypeDirectionDescription
inputconst int8_t *inPointer to the input int8_t array.
outputfloat *outPointer to the output float array.
sizeint32_tinNumber of elements in the arrays.
zero_pointint32_tinZero point (offset) that was used during quantization.
scalefloatinScale factor that was used during quantization.
Returns of arm_dequantize_s8_f32
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Dequantize an int16t array back to floating-point format.

Include/arm_nnfunctions.h:7442

arm_cmsis_nn_status arm_dequantize_s16_f32(
const int16_t *input,
float *output,
int32_t size,
int32_t zero_point,
float scale
)

Dequantize an int16_t array back to floating-point format.

Parameters of arm_dequantize_s16_f32
NameTypeDirectionDescription
inputconst int16_t *inPointer to the input int16_t array.
outputfloat *outPointer to the output float array.
sizeint32_tinNumber of elements in the arrays.
zero_pointint32_tinZero point (offset) that was used during quantization.
scalefloatinScale factor that was used during quantization.
Returns of arm_dequantize_s16_f32
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Strided slice function for int8 data.

Include/arm_nnfunctions.h:7465

arm_cmsis_nn_status arm_strided_slice_s8(
const int8_t *input_data,
int8_t *output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const begin_dims,
const cmsis_nn_dims *const stride_dims,
const cmsis_nn_dims *const output_dims
)

Strided slice function for int8 data.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_strided_slice_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to input tensor
output_dataint8_t *outPointer to output tensor
input_dimsconst cmsis_nn_dims *constinInput tensor dimensions
begin_dimsconst cmsis_nn_dims *constinBegin dimensions for slicing
stride_dimsconst cmsis_nn_dims *constinStride dimensions for slicing
output_dimsconst cmsis_nn_dims *constinOutput tensor dimensions
Returns of arm_strided_slice_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Strided slice function for int16 data.

Include/arm_nnfunctions.h:7488

arm_cmsis_nn_status arm_strided_slice_s16(
const int16_t *input_data,
int16_t *output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const begin_dims,
const cmsis_nn_dims *const stride_dims,
const cmsis_nn_dims *const output_dims
)

Strided slice function for int16 data.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_strided_slice_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to input tensor
output_dataint16_t *outPointer to output tensor
input_dimsconst cmsis_nn_dims *constinInput tensor dimensions
begin_dimsconst cmsis_nn_dims *constinBegin dimensions for slicing
stride_dimsconst cmsis_nn_dims *constinStride dimensions for slicing
output_dimsconst cmsis_nn_dims *constinOutput tensor dimensions
Returns of arm_strided_slice_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Strided slice function for int32 data.

Include/arm_nnfunctions.h:7511

arm_cmsis_nn_status arm_strided_slice_s32(
const int32_t *input_data,
int32_t *output_data,
const cmsis_nn_dims *const input_dims,
const cmsis_nn_dims *const begin_dims,
const cmsis_nn_dims *const stride_dims,
const cmsis_nn_dims *const output_dims
)

Strided slice function for int32 data.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_strided_slice_s32
NameTypeDirectionDescription
input_dataconst int32_t *inPointer to input tensor
output_dataint32_t *outPointer to output tensor
input_dimsconst cmsis_nn_dims *constinInput tensor dimensions
begin_dimsconst cmsis_nn_dims *constinBegin dimensions for slicing
stride_dimsconst cmsis_nn_dims *constinStride dimensions for slicing
output_dimsconst cmsis_nn_dims *constinOutput tensor dimensions
Returns of arm_strided_slice_s32
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Gather elements along an axis for int8 tensors.

Include/arm_nnfunctions.h:7540

arm_cmsis_nn_status arm_gather_s8(
const int8_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_params *params,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)

Gather elements along an axis for int8 tensors.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_gather_s8
NameTypeDirectionDescription
input_dataconst int8_t *inPointer to input tensor data
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
indices_dataconst int32_t *inPointer to indices tensor data (int32)
indices_dimsconst cmsis_nn_dims *inIndices tensor dimensions
paramsconst cmsis_nn_gather_params *inPointer to gather parameters
output_dataint8_t *outPointer to output tensor data
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_gather_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Gather elements along an axis for int16 tensors.

Include/arm_nnfunctions.h:7565

arm_cmsis_nn_status arm_gather_s16(
const int16_t *input_data,
const cmsis_nn_dims *input_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_params *params,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)

Gather elements along an axis for int16 tensors.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_gather_s16
NameTypeDirectionDescription
input_dataconst int16_t *inPointer to input tensor data
input_dimsconst cmsis_nn_dims *inInput tensor dimensions
indices_dataconst int32_t *inPointer to indices tensor data (int32)
indices_dimsconst cmsis_nn_dims *inIndices tensor dimensions
paramsconst cmsis_nn_gather_params *inPointer to gather parameters
output_dataint16_t *outPointer to output tensor data
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_gather_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Gathernd slices for int8 tensors.

Include/arm_nnfunctions.h:7590

arm_cmsis_nn_status arm_gather_nd_s8(
const int8_t *params_data,
const cmsis_nn_dims *params_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_nd_params *params,
int8_t *output_data,
const cmsis_nn_dims *output_dims
)

Gather_nd slices for int8 tensors.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_gather_nd_s8
NameTypeDirectionDescription
params_dataconst int8_t *inPointer to params tensor data
params_dimsconst cmsis_nn_dims *inParams tensor dimensions
indices_dataconst int32_t *inPointer to indices tensor data (int32)
indices_dimsconst cmsis_nn_dims *inIndices tensor dimensions
paramsconst cmsis_nn_gather_nd_params *inPointer to gather_nd parameters
output_dataint8_t *outPointer to output tensor data
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_gather_nd_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Gathernd slices for int16 tensors.

Include/arm_nnfunctions.h:7615

arm_cmsis_nn_status arm_gather_nd_s16(
const int16_t *params_data,
const cmsis_nn_dims *params_dims,
const int32_t *indices_data,
const cmsis_nn_dims *indices_dims,
const cmsis_nn_gather_nd_params *params,
int16_t *output_data,
const cmsis_nn_dims *output_dims
)

Gather_nd slices for int16 tensors.

  1. Supported framework: TensorFlow Lite Micro
Parameters of arm_gather_nd_s16
NameTypeDirectionDescription
params_dataconst int16_t *inPointer to params tensor data
params_dimsconst cmsis_nn_dims *inParams tensor dimensions
indices_dataconst int32_t *inPointer to indices tensor data (int32)
indices_dimsconst cmsis_nn_dims *inIndices tensor dimensions
paramsconst cmsis_nn_gather_nd_params *inPointer to gather_nd parameters
output_dataint16_t *outPointer to output tensor data
output_dimsconst cmsis_nn_dims *inOutput tensor dimensions
Returns of arm_gather_nd_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Tile an int8 tensor along each dimension.

Include/arm_nnfunctions.h:7642

arm_cmsis_nn_status arm_tile_s8(const int8_t *input, const cmsis_nn_tile_params *params, int8_t *output)

Tile an int8 tensor along each dimension.

  1. Supported framework: TensorFlow Lite Micro
  2. Maximum rank: 8
Parameters of arm_tile_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to input tensor data
paramsconst cmsis_nn_tile_params *inPointer to tile parameters (rank, input_shape, multiples)
outputint8_t *outPointer to output tensor data (pre-allocated by caller)
Returns of arm_tile_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Tile an int16 tensor along each dimension.

Include/arm_nnfunctions.h:7654

arm_cmsis_nn_status arm_tile_s16(const int16_t *input, const cmsis_nn_tile_params *params, int16_t *output)

Tile an int16 tensor along each dimension.

Parameters of arm_tile_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to input tensor data
paramsconst cmsis_nn_tile_params *inPointer to tile parameters (rank, input_shape, multiples)
outputint16_t *outPointer to output tensor data (pre-allocated by caller)
Returns of arm_tile_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Broadcast an int8 tensor to a target shape.

Include/arm_nnfunctions.h:7676

arm_cmsis_nn_status arm_broadcast_to_s8(const int8_t *input, const cmsis_nn_broadcast_to_params *params, int8_t *output)

Broadcast an int8 tensor to a target shape.

  1. Input dimensions must be 1 or match the output dimension for each axis.
  2. Maximum rank: 8
Parameters of arm_broadcast_to_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to input tensor data
paramsconst cmsis_nn_broadcast_to_params *inPointer to broadcast parameters (rank, input/output shapes)
outputint8_t *outPointer to output tensor data (pre-allocated by caller)
Returns of arm_broadcast_to_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Broadcast an int16 tensor to a target shape.

Include/arm_nnfunctions.h:7689

arm_cmsis_nn_status arm_broadcast_to_s16(
const int16_t *input,
const cmsis_nn_broadcast_to_params *params,
int16_t *output
)

Broadcast an int16 tensor to a target shape.

Parameters of arm_broadcast_to_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to input tensor data
paramsconst cmsis_nn_broadcast_to_params *inPointer to broadcast parameters (rank, input/output shapes)
outputint16_t *outPointer to output tensor data (pre-allocated by caller)
Returns of arm_broadcast_to_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Scatter updates into a zero-initialized output tensor for int8.

Include/arm_nnfunctions.h:7707

arm_cmsis_nn_status arm_scatter_nd_s8(
const int32_t *indices,
const int8_t *updates,
const cmsis_nn_scatter_nd_params *params,
int8_t *output
)

Scatter updates into a zero-initialized output tensor for int8.

Parameters of arm_scatter_nd_s8
NameTypeDirectionDescription
indicesconst int32_t *inPointer to indices data (int32, shape [num_updates, index_depth])
updatesconst int8_t *inPointer to updates data
paramsconst cmsis_nn_scatter_nd_params *inPointer to scatter_nd parameters
outputint8_t *outPointer to output tensor data (pre-allocated by caller)
Returns of arm_scatter_nd_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Scatter updates into a zero-initialized output tensor for int16.

Include/arm_nnfunctions.h:7723

arm_cmsis_nn_status arm_scatter_nd_s16(
const int32_t *indices,
const int16_t *updates,
const cmsis_nn_scatter_nd_params *params,
int16_t *output
)

Scatter updates into a zero-initialized output tensor for int16.

Parameters of arm_scatter_nd_s16
NameTypeDirectionDescription
indicesconst int32_t *inPointer to indices data (int32, shape [num_updates, index_depth])
updatesconst int16_t *inPointer to updates data
paramsconst cmsis_nn_scatter_nd_params *inPointer to scatter_nd parameters
outputint16_t *outPointer to output tensor data (pre-allocated by caller)
Returns of arm_scatter_nd_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode).

Include/arm_nnfunctions.h:7743

arm_cmsis_nn_status arm_mirror_pad_s8(const int8_t *input, const cmsis_nn_mirror_pad_params *params, int8_t *output)

Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode).

Parameters of arm_mirror_pad_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to input tensor data
paramsconst cmsis_nn_mirror_pad_params *inPointer to mirror_pad parameters
outputint8_t *outPointer to output tensor data (pre-allocated by caller)
Returns of arm_mirror_pad_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode).

Include/arm_nnfunctions.h:7755

arm_cmsis_nn_status arm_mirror_pad_s16(const int16_t *input, const cmsis_nn_mirror_pad_params *params, int16_t *output)

Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode).

Parameters of arm_mirror_pad_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to input tensor data
paramsconst cmsis_nn_mirror_pad_params *inPointer to mirror_pad parameters
outputint16_t *outPointer to output tensor data (pre-allocated by caller)
Returns of arm_mirror_pad_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

WHERE operator: return coordinates of non-zero elements in condition.

Include/arm_nnfunctions.h:7774

arm_cmsis_nn_status arm_where_s8(
const int8_t *condition,
const cmsis_nn_where_params *params,
int64_t *output,
int32_t *num_true
)

WHERE operator: return coordinates of non-zero elements in condition.

Parameters of arm_where_s8
NameTypeDirectionDescription
conditionconst int8_t *inPointer to condition tensor data (int8, non-zero = true)
paramsconst cmsis_nn_where_params *inPointer to where parameters (rank, shape)
outputint64_t *outPointer to output coordinates (int64, shape [max_true, rank])
num_trueint32_t *outNumber of true elements found
Returns of arm_where_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

WHERE operator: return coordinates of non-zero elements in condition (int16).

Include/arm_nnfunctions.h:7788

arm_cmsis_nn_status arm_where_s16(
const int16_t *condition,
const cmsis_nn_where_params *params,
int64_t *output,
int32_t *num_true
)

WHERE operator: return coordinates of non-zero elements in condition (int16).

Parameters of arm_where_s16
NameTypeDirectionDescription
conditionconst int16_t *inPointer to condition tensor data (int16, non-zero = true)
paramsconst cmsis_nn_where_params *inPointer to where parameters (rank, shape)
outputint64_t *outPointer to output coordinates (int64, shape [max_true, rank])
num_trueint32_t *outNumber of true elements found
Returns of arm_where_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

SELECTV2 with broadcast for int8 tensors.

Include/arm_nnfunctions.h:7802

arm_cmsis_nn_status arm_select_v2_s8(
const bool *condition,
const int8_t *x,
const int8_t *y,
const cmsis_nn_select_v2_params *params,
int8_t *output
)

SELECT_V2 with broadcast for int8 tensors.

Parameters of arm_select_v2_s8
NameTypeDirectionDescription
conditionconst bool *inPointer to condition tensor data (bool)
xconst int8_t *inPointer to x tensor data (selected when condition is true)
yconst int8_t *inPointer to y tensor data (selected when condition is false)
paramsconst cmsis_nn_select_v2_params *inPointer to select_v2 parameters (broadcast strides)
outputint8_t *outPointer to output tensor data
Returns of arm_select_v2_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

SELECTV2 with broadcast for int16 tensors.

Include/arm_nnfunctions.h:7820

arm_cmsis_nn_status arm_select_v2_s16(
const bool *condition,
const int16_t *x,
const int16_t *y,
const cmsis_nn_select_v2_params *params,
int16_t *output
)

SELECT_V2 with broadcast for int16 tensors.

Parameters of arm_select_v2_s16
NameTypeDirectionDescription
conditionconst bool *inPointer to condition tensor data (bool)
xconst int16_t *inPointer to x tensor data
yconst int16_t *inPointer to y tensor data
paramsconst cmsis_nn_select_v2_params *inPointer to select_v2 parameters (broadcast strides)
outputint16_t *outPointer to output tensor data
Returns of arm_select_v2_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Reverse variable-length sequences along a dimension for int8.

Include/arm_nnfunctions.h:7842

arm_cmsis_nn_status arm_reverse_sequence_s8(
const int8_t *input,
const int32_t *seq_lengths,
const cmsis_nn_reverse_sequence_params *params,
int8_t *output
)

Reverse variable-length sequences along a dimension for int8.

Parameters of arm_reverse_sequence_s8
NameTypeDirectionDescription
inputconst int8_t *inPointer to input tensor data
seq_lengthsconst int32_t *inPointer to per-batch sequence lengths (int32)
paramsconst cmsis_nn_reverse_sequence_params *inPointer to reverse_sequence parameters
outputint8_t *outPointer to output tensor data
Returns of arm_reverse_sequence_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Reverse variable-length sequences along a dimension for int16.

Include/arm_nnfunctions.h:7858

arm_cmsis_nn_status arm_reverse_sequence_s16(
const int16_t *input,
const int32_t *seq_lengths,
const cmsis_nn_reverse_sequence_params *params,
int16_t *output
)

Reverse variable-length sequences along a dimension for int16.

Parameters of arm_reverse_sequence_s16
NameTypeDirectionDescription
inputconst int16_t *inPointer to input tensor data
seq_lengthsconst int32_t *inPointer to per-batch sequence lengths (int32)
paramsconst cmsis_nn_reverse_sequence_params *inPointer to reverse_sequence parameters
outputint16_t *outPointer to output tensor data
Returns of arm_reverse_sequence_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Update a slice of an int8 operand tensor at runtime-determined indices.

Include/arm_nnfunctions.h:7880

arm_cmsis_nn_status arm_dynamic_update_slice_s8(
const int8_t *operand,
const int8_t *update,
const int32_t *start_indices,
const cmsis_nn_dynamic_update_slice_params *params,
int8_t *output
)

Update a slice of an int8 operand tensor at runtime-determined indices.

Parameters of arm_dynamic_update_slice_s8
NameTypeDirectionDescription
operandconst int8_t *inPointer to operand tensor data (copied to output first)
updateconst int8_t *inPointer to update tensor data
start_indicesconst int32_t *inPointer to start index per dimension (int32, length = rank)
paramsconst cmsis_nn_dynamic_update_slice_params *inPointer to dynamic_update_slice parameters
outputint8_t *outPointer to output tensor data
Returns of arm_dynamic_update_slice_s8
Description
The function returns `ARM_CMSIS_NN_SUCCESS`
function

Update a slice of an int16 operand tensor at runtime-determined indices.

Include/arm_nnfunctions.h:7898

arm_cmsis_nn_status arm_dynamic_update_slice_s16(
const int16_t *operand,
const int16_t *update,
const int32_t *start_indices,
const cmsis_nn_dynamic_update_slice_params *params,
int16_t *output
)

Update a slice of an int16 operand tensor at runtime-determined indices.

Parameters of arm_dynamic_update_slice_s16
NameTypeDirectionDescription
operandconst int16_t *inPointer to operand tensor data (copied to output first)
updateconst int16_t *inPointer to update tensor data
start_indicesconst int32_t *inPointer to start index per dimension (int32, length = rank)
paramsconst cmsis_nn_dynamic_update_slice_params *inPointer to dynamic_update_slice parameters
outputint16_t *outPointer to output tensor data
Returns of arm_dynamic_update_slice_s16
Description
The function returns `ARM_CMSIS_NN_SUCCESS`