#define USE_INTRINSICarm_nnfunctions.h
USE_INTRINSICmacroarm_convolve_wrapper_s4functions4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.arm_convolve_wrapper_s4_get_buffer_sizefunctionGet the required buffer size for armconvolvewrappers4.arm_convolve_wrapper_s4_get_buffer_size_mvefunctionGet the required buffer size for armconvolvewrappers4 for Arm(R) Helium Architecture case.arm_convolve_wrapper_s4_get_buffer_size_dspfunctionGet the required buffer size for armconvolvewrappers4 for processors with DSP extension.arm_convolve_wrapper_s8functions8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.arm_convolve_wrapper_s8_get_buffer_sizefunctionGet the required buffer size for armconvolvewrappers8.arm_convolve_s8_get_buffer_size_mvefunctionGet the required buffer size for armconvolves8 for Arm(R) Helium Architecture case.arm_convolve_wrapper_s8_get_buffer_size_mvefunctionGet the required buffer size for armconvolvewrappers8 for Arm(R) Helium Architecture case.arm_convolve_wrapper_s8_get_buffer_size_dspfunctionGet the required buffer size for armconvolvewrappers8 for processors with DSP extension.arm_convolve_wrapper_s16functions16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.arm_convolve_s16_group_ch_mult_1functions16 grouped convolution optimized for the case where filterdims->c == 1 and inputch == outputch (channel multiplier = 1).arm_convolve_wrapper_s16_get_buffer_sizefunctionGet the required buffer size for armconvolvewrappers16.arm_convolve_wrapper_s16_get_buffer_size_dspfunctionGet the required buffer size for armconvolvewrappers16 for for processors with DSP extension.arm_convolve_wrapper_s16_get_buffer_size_mvefunctionGet the required buffer size for armconvolvewrappers16 for Arm(R) Helium Architecture case.arm_convolve_s4functionBasic s4 convolution function.arm_convolve_even_s4functionBasic s4 convolution function with a requirement of even number of kernels.arm_convolve_s8functionBasic s8 convolution function.arm_convolve_s8_small_cinfunctions8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model.arm_convolve_s8_3x3_c16_s1functions8 3x3 convolution over 16 input channels with unit stride.arm_convolve_s4_get_buffer_sizefunctionGet the required buffer size for s4 convolution function.arm_convolve_even_s4_get_buffer_sizefunctionGet the required buffer size for armconvolveevens4.arm_convolve_s8_get_buffer_sizefunctionGet the required buffer size for s8 convolution function.arm_convolve_s8_get_weights_sum_sizefunctionGet the required buffer size for s8 convolution and depthwise convolution weight sum.arm_transpose_conv_wrapper_s8functionWrapper to select optimal transposed convolution algorithm depending on parameters.arm_transpose_conv_s8functionBasic s8 transpose convolution function.arm_transpose_conv_s8_get_buffer_sizefunctionGet the required buffer size for ctx in s8 transpose conv function.arm_transpose_conv_s8_get_reverse_conv_buffer_sizefunctionGet the required buffer size for outputctx in s8 transpose conv function.arm_transpose_conv_s8_get_buffer_size_mvefunctionGet size of additional buffer required by armtransposeconvs8() for Arm(R) Helium Architecture case.arm_convolve_s16functionBasic s16 convolution function.arm_convolve_1x1_s16_ns_np_ndfunctionPointwise s16 convolution function: no stride, no padding, no dilation.arm_convolve_s16_fast_small_kernelfunctionarmconvolves16fastsmallkernel function.arm_convolve_s16_get_buffer_sizefunctionGet the required buffer size for s16 convolution function.arm_convolve_1x1_s4_fastfunctionFast s4 version for 1x1 convolution (non-square shape).arm_convolve_1x1_s4functions4 version for 1x1 convolution with support for non-unity stride valuesarm_convolve_1x1_s8_fastfunctionFast s8 version for 1x1 convolution (non-square shape).arm_convolve_1x1_s4_fast_get_buffer_sizefunctionGet the required buffer size for armconvolve1x1s4fast.arm_convolve_1x1_s8_fast_get_buffer_sizefunctionGet the required buffer size for armconvolve1x1s8fast.arm_convolve_1x1_s8functions8 version for 1x1 convolution with support for non-unity stride valuesarm_convolve_1_x_n_s8function1xn convolutionarm_convolve_weight_sumfunctionPre-computes per-output-channel weight sums for a standard convolution.arm_depthwise_convolve_weight_sumfunctionPre-computes per-channel weight sums for a depthwise convolution.arm_convolve_1x1_out_s8functionOptimised convolution for 1x1 output images (shape of BX1x1xCOUT) for 8x8 computations.arm_convolve_1x1_out_s8_get_buffer_sizefunctionGet the required scratch buffer size for armconvolve1x1outs8().arm_convolve_1_x_n_s4function1xn convolution for s4 weightsarm_convolve_1_x_n_s8_get_buffer_sizefunctionGet the required additional buffer size for 1xn convolution.arm_convolve_1_x_n_s4_get_buffer_sizefunctionGet the required additional buffer size for 1xn convolution.arm_depthwise_conv_wrapper_s8functionWrapper function to pick the right optimized s8 depthwise convolution function.arm_depthwise_conv_wrapper_s4functionWrapper function to pick the right optimized s4 depthwise convolution function.arm_depthwise_conv_wrapper_s8_get_buffer_sizefunctionGet size of additional buffer required by armdepthwiseconvwrappers8().arm_depthwise_conv_wrapper_s8_get_buffer_size_dspfunctionGet size of additional buffer required by armdepthwiseconvwrappers8() for processors with DSP extension.arm_depthwise_conv_wrapper_s8_get_buffer_size_mvefunctionGet size of additional buffer required by armdepthwiseconvwrappers8() for Arm(R) Helium Architecture case.arm_depthwise_conv_wrapper_s4_get_buffer_sizefunctionGet size of additional buffer required by armdepthwiseconvwrappers4().arm_depthwise_conv_wrapper_s4_get_buffer_size_dspfunctionGet size of additional buffer required by armdepthwiseconvwrappers4() for processors with DSP extension.arm_depthwise_conv_wrapper_s4_get_buffer_size_mvefunctionGet size of additional buffer required by armdepthwiseconvwrappers4() for Arm(R) Helium Architecture case.arm_depthwise_conv_s8functionBasic s8 depthwise convolution function that doesn't have any constraints on the input dimensions.arm_depthwise_conv_s4functionBasic s4 depthwise convolution function that doesn't have any constraints on the input dimensions.arm_depthwise_conv_s16functionBasic s16 depthwise convolution function that doesn't have any constraints on the input dimensions.arm_depthwise_conv_wrapper_s16functionWrapper function to pick the right optimized s16 depthwise convolution function.arm_depthwise_conv_wrapper_s16_get_buffer_sizefunctionGet size of additional buffer required by armdepthwiseconvwrappers16().arm_depthwise_conv_wrapper_s16_get_buffer_size_dspfunctionGet size of additional buffer required by armdepthwiseconvwrappers16() for processors with DSP extension.arm_depthwise_conv_wrapper_s16_get_buffer_size_mvefunctionGet size of additional buffer required by armdepthwiseconvwrappers16() for Arm(R) Helium Architecture case.arm_depthwise_conv_fast_s16functionOptimized s16 depthwise convolution function with constraint that inchannel equals outchannel.arm_depthwise_conv_fast_s16_get_buffer_sizefunctionGet the required buffer size for optimized s16 depthwise convolution function with constraint that inchannel equals outchannel.arm_depthwise_conv_3x3_s8functionOptimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below).arm_depthwise_conv_s8_optfunctionOptimized s8 depthwise convolution function with constraint that inchannel equals outchannel.arm_depthwise_conv_s8_opt_planar_supportedfunctionWhether armdepthwiseconvs8opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels.arm_depthwise_conv_s8_opt_planarfunctionThe planar path of armdepthwiseconvs8opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane.arm_depthwise_conv_s8_opt_channelwisefunctionThe channel-vectorized path of armdepthwiseconvs8opt() on its own, without the planar attempt.arm_depthwise_conv_s8_opt_3x3functions8 3x3 depthwise convolution that reads the input in place, without the im2col copy of armdepthwiseconvs8opt(), for layers in its gate.arm_depthwise_conv_s8_opt_3x3_c64_s1functionarmdepthwiseconvs8opt3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps.arm_depthwise_conv_s8_opt_3x3_get_buffer_sizefunctionGet the scratch size in bytes of armdepthwiseconvs8opt3x3() and armdepthwiseconvs8opt3x3c64s1().arm_depthwise_conv_s4_optfunctionOptimized s4 depthwise convolution function with constraint that inchannel equals outchannel.arm_depthwise_conv_s8_opt_get_buffer_sizefunctionGet the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel.arm_depthwise_conv_s4_opt_get_buffer_sizefunctionGet the required buffer size for optimized s4 depthwise convolution function with constraint that inchannel equals outchannel.arm_fully_connected_s4functionBasic s4 Fully Connected function.arm_fully_connected_s8functionBasic s8 Fully Connected function.arm_fully_connected_per_channel_s8functionBasic s8 Fully Connected function using per channel quantization.arm_fully_connected_wrapper_s8functions8 Fully Connected layer wrapper functionarm_vector_sum_s8functionCalculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s32 biasdata.arm_vector_sum_s8_s64functionCalculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s64 biasdata.arm_fully_connected_s8_get_buffer_sizefunctionGet size of additional buffer required by armfullyconnecteds8().arm_fully_connected_s8_get_buffer_size_dspfunctionGet size of additional buffer required by armfullyconnecteds8() for processors with DSP extension.arm_fully_connected_s8_get_buffer_size_mvefunctionGet size of additional buffer required by armfullyconnecteds8() for Arm(R) Helium Architecture case.arm_fully_connected_s16functionBasic s16 Fully Connected function.arm_fully_connected_per_channel_s16functionBasic s16 Fully Connected function using per channel quantization.arm_fully_connected_wrapper_s16functions16 Fully Connected layer wrapper functionarm_fully_connected_s16_get_buffer_sizefunctionGet size of additional buffer required by armfullyconnecteds16().arm_fully_connected_s16_get_buffer_size_dspfunctionGet size of additional buffer required by armfullyconnecteds16() for processors with DSP extension.arm_fully_connected_s16_get_buffer_size_mvefunctionGet size of additional buffer required by armfullyconnecteds16() for Arm(R) Helium Architecture case.arm_fully_connected_per_channel_s16_get_buffer_sizefunctionGet size of additional buffer required by armfullyconnectedperchannels16().arm_fully_connected_per_channel_s16_get_buffer_size_dspfunctionGet size of additional buffer required by armfullyconnectedperchannels16() for processors with DSP extension.arm_fully_connected_per_channel_s16_get_buffer_size_mvefunctionGet size of additional buffer required by armfullyconnectedperchannels16() for Arm(R) Helium Architecture case.arm_add_s8functions8 elementwise add of two tensors with support for broadcasting.arm_add_scalar_s8functions8 elementwise add of scalar and vectorarm_elementwise_add_s8functions8 elementwise add of two vectorsarm_abs_s8functions8 elementwise absolute valuearm_sqrt_s8functions8 elementwise square rootarm_sqrt_s16functions16 elementwise square root using piecewise LUT with linear interpolationarm_sqrt_s16_tablefreefunctions16 elementwise square root without a lookup tablearm_abs_s16functions16 elementwise absolute valuearm_rsqrt_s16_per_opfunctionINT16 reciprocal square root using a per-operator LUT.arm_rsqrt_s16_universalfunctionINT16 reciprocal square root using a shared universal LUT.arm_sub_s8functions8 elementwise subtraction of two tensors with support for broadcasting.arm_sub_scalar_s8functions8 elementwise subtract of scalar and vector (scalar - vector)arm_elementwise_sub_s8functions8 elementwise subtract of two vectorsarm_add_s16functions16 elementwise add of two tensors with support for broadcasting.arm_add_scalar_s16functions16 elementwise add of scalar and vectorarm_elementwise_add_s16functions16 elementwise add of two vectorsarm_sub_s16functions16 elementwise subtraction of two tensors with support for broadcasting.arm_sub_scalar_s16functions16 elementwise subtract of scalar and vector (scalar - vector)arm_elementwise_sub_s16functions16 elementwise subtract of two vectorsarm_squared_difference_s8functions8 elementwise squared difference of two tensors with support for broadcasting.arm_squared_difference_scalar_s8functions8 elementwise squared difference of scalar and vector.arm_elementwise_squared_difference_s8functions8 elementwise squared difference of two vectors.arm_squared_difference_s16functions16 elementwise squared difference of two tensors with support for broadcasting.arm_squared_difference_scalar_s16functions16 elementwise squared difference of scalar and vector.arm_elementwise_squared_difference_s16functions16 elementwise squared difference of two vectors.arm_mul_s8functions8 elementwise multiplication of two tensors with support for broadcasting.arm_mul_scalar_s8functions8 elementwise multiplication of scalar and vectorarm_elementwise_mul_s8functions8 elementwise multiplicationarm_mul_s16functions16 elementwise multiplication of two tensors with support for broadcasting.arm_mul_scalar_s16functions16 elementwise multiplication of scalar and vectorarm_elementwise_mul_s16functions16 elementwise multiplicationarm_minimum_s8functions8 elementwise minimum w/ support for broadcasting and scalar inputs.arm_maximum_s8functions8 elementwise maximum w/ support for broadcasting and scalar inputs.arm_minimum_s16functions16 elementwise minimum w/ support for broadcasting and scalar inputs.arm_maximum_s16functions16 elementwise maximum w/ support for broadcasting and scalar inputs.arm_comparison_s8functions8 elementwise comparison with support for broadcasting.arm_comparison_s16functions16 elementwise comparison with support for broadcasting.arm_equal_s8functions8 elementwise equality comparison with support for broadcasting.arm_not_equal_s8functions8 elementwise inequality comparison with support for broadcasting.arm_greater_s8functions8 elementwise greater-than comparison with support for broadcasting.arm_greater_equal_s8functions8 elementwise greater-or-equal comparison with support for broadcasting.arm_less_s8functions8 elementwise less-than comparison with support for broadcasting.arm_less_equal_s8functions8 elementwise less-or-equal comparison with support for broadcasting.arm_equal_s16functions16 elementwise equality comparison with support for broadcasting.arm_not_equal_s16functions16 elementwise inequality comparison with support for broadcasting.arm_greater_s16functions16 elementwise greater-than comparison with support for broadcasting.arm_greater_equal_s16functions16 elementwise greater-or-equal comparison with support for broadcasting.arm_less_s16functions16 elementwise less-than comparison with support for broadcasting.arm_less_equal_s16functions16 elementwise less-or-equal comparison with support for broadcasting.arm_relu_q7functionQ7 RELU function.arm_relu6_q7functionQ7 RELU6 function.arm_relu_q15functionQ15 RELU function.arm_clamp_s8functionS8 clamp function.arm_clamp_s16functionS16 clamp function.arm_relu_s8functionS8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127.arm_relu_generic_s8functionS8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.arm_relu_s16functionS16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767.arm_relu_generic_s16functionS16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.arm_leaky_relu_s8functionS8 Leaky ReLU activation function.arm_leaky_relu_s16functionS16 Leaky ReLU activation function.arm_logistic_s16functionLogistic activation function for s16.arm_tanh_s16functionTanh activation function for s16.arm_nn_activation_s16functions16 neural network activation function using direct table look-uparm_hard_swish_compat_s8functionS8 Hard-Swish activation function (compatibility version).arm_hard_swish_precise_s8functionS8 Hard-Swish activation function (precise version).arm_hard_swish_precise_s16functionS16 Hard-Swish activation function (precise version).arm_prelu_s8functionS8 PReLU activation function.arm_elementwise_prelu_s8functionElementwise S8 PReLU activation function.arm_prelu_scalar_s8functionScalar S8 PReLU activation function.arm_prelu_s16functionS16 PReLU activation function.arm_elementwise_prelu_s16functionElementwise S16 PReLU activation function.arm_prelu_scalar_s16functionScalar S16 PReLU activation function.arm_avgpool_s8functions8 average pooling function.arm_avgpool_s8_get_buffer_sizefunctionGet the required buffer size for S8 average pooling function.arm_avgpool_s8_get_buffer_size_dspfunctionGet the required buffer size for S8 average pooling function for processors with DSP extension.arm_avgpool_s8_get_buffer_size_mvefunctionGet the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case.arm_avgpool_s16functions16 average pooling function.arm_avgpool_s16_get_buffer_sizefunctionGet the required buffer size for S16 average pooling function.arm_avgpool_s16_get_buffer_size_dspfunctionGet the required buffer size for S16 average pooling function for processors with DSP extension.arm_avgpool_s16_get_buffer_size_mvefunctionGet the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case.arm_max_pool_s8functions8 max pooling function.arm_max_pool_s16functions16 max pooling function.arm_softmax_s8functionS8 softmax function.arm_softmax_s8_s16functionS8 to s16 softmax function.arm_softmax_s16functionS16 softmax function.arm_softmax_u8functionU8 softmax function.arm_reshape_s8functionReshape a s8 vector into another with different shape.arm_resize_nearest_neighbor_s8functionNearest neighbor resize function for s8 data.arm_resize_nearest_neighbor_s16functionNearest neighbor resize function for s16 data.arm_space_to_depth_s8functionSpace to Depth function for s8 data type.arm_space_to_depth_s16functionSpace to Depth function for s16 data type.arm_depth_to_space_s8functionDepth to Space function for s8 data type.arm_depth_to_space_s16functionDepth to Space function for s16 data type.arm_space_to_batch_nd_s8functionSpace to Batch ND function for s8 data type.arm_space_to_batch_nd_s16functionSpace to Batch ND function for s16 data type.arm_batch_to_space_nd_s8functionBatch to Space ND function for s8 data type.arm_batch_to_space_nd_s16functionBatch to Space ND function for s16 data type.arm_transpose_s8functionBasic transpose function.arm_transpose_s16functionBasic s16 transpose function.arm_concatenation_s8_xfunctionint8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate.arm_concatenation_s8_yfunctionint8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate.arm_concatenation_s8_zfunctionint8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate.arm_concatenation_s8_wfunctionint8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to…arm_concatenation_s8functionint8/uint8 concatenation function to be used for concatenating N-tensors along the target axisarm_concatenation_s16functionint16/uint16 concatenation function to be used for concatenating N-tensors along the target axisarm_concatenation_s32functionint32/uint32 concatenation function to be used for concatenating N-tensors along the target axisarm_split_s8functionint8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axisarm_split_s16functionint16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axisarm_svdf_s8functions8 SVDF function with 8 bit state tensor and 8 bit time weightsarm_svdf_state_s16_s8functions8 SVDF function with 16 bit state tensor and 16 bit time weightsarm_svdf_s8_get_buffer_sizefunctionGet size of the kernel-sum buffer required by armsvdfs8().arm_svdf_s8_get_buffer_size_dspfunctionGet size of the kernel-sum buffer required by armsvdfs8() for processors with DSP extension.arm_svdf_s8_get_buffer_size_mvefunctionGet size of the kernel-sum buffer required by armsvdfs8() for Arm(R) Helium Architecture case.arm_svdf_s8_input_ctx_get_buffer_sizefunctionGet size of the inputctx staging buffer required by armsvdfs8().arm_svdf_s8_output_ctx_get_buffer_sizefunctionGet size of the outputctx staging buffer required by armsvdfs8().arm_svdf_state_s16_s8_input_ctx_get_buffer_sizefunctionGet size of the inputctx staging buffer required by armsvdfstates16s8().arm_svdf_state_s16_s8_output_ctx_get_buffer_sizefunctionGet size of the outputctx staging buffer required by armsvdfstates16s8().arm_lstm_unidirectional_s8functionLSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias.arm_lstm_unidirectional_s16functionLSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias.arm_lstm_unidirectional_s8_temp1_get_buffer_sizefunctionGet size of the temp1 scratch buffer required by armlstmunidirectionals8().arm_lstm_unidirectional_s8_temp2_get_buffer_sizefunctionGet size of the temp2 scratch buffer required by armlstmunidirectionals8().arm_lstm_unidirectional_s16_temp1_get_buffer_sizefunctionGet size of the temp1 scratch buffer required by armlstmunidirectionals16().arm_lstm_unidirectional_s16_temp2_get_buffer_sizefunctionGet size of the temp2 scratch buffer required by armlstmunidirectionals16().arm_batch_matmul_s8functionBatch matmul function with 8 bit input and output.arm_batch_matmul_s16functionBatch matmul function with 16 bit input and output.arm_batch_matmul_s8_get_buffer_sizefunctionGet size of the scratch buffer required by armbatchmatmuls8().arm_batch_matmul_s8_get_buffer_size_dspfunctionGet size of the scratch buffer required by armbatchmatmuls8() for processors with DSP extension.arm_batch_matmul_s8_get_buffer_size_mvefunctionGet size of the scratch buffer required by armbatchmatmuls8() for Arm(R) Helium Architecture case.arm_pad_s8functionExpands the size of the input by adding constant values before and after the data, in all dimensions.arm_pad_s16functionExpands the size of the input by adding constant values before and after the data, in all dimensions.arm_mean_s8functionComputes the mean of the input tensor along the specified axis.arm_mean_s16functionComputes the mean of the input tensor along the specified axis.arm_argmax_s8functionCompute ArgMax indices of an s8 tensor along a specific axis.arm_argmin_s8functionCompute ArgMin indices of an s8 tensor along a specific axis.arm_argmax_s16functionCompute ArgMax indices of an s16 tensor along a specific axis.arm_argmin_s16functionCompute ArgMin indices of an s16 tensor along a specific axis.arm_reduce_max_s8functionComputes the max of the input tensor along the specified axis.arm_reduce_max_s16functionComputes the max of the input tensor along the specified axis.arm_reduce_min_s8functionComputes the min of the input tensor along the specified axis.arm_reduce_min_s16functionComputes the min of the input tensor along the specified axis.arm_quantize_f32_s8functionQuantize a floating-point array into int8t format.arm_quantize_f32_s16functionQuantize a floating-point array into int16t format.arm_requantize_s8_s8functionRequantize an int8t array to another int8t range with a different scale.arm_requantize_s16_s16functionRequantize an int16t array to another int16t range with a different scale.arm_dequantize_s8_f32functionDequantize an int8t array back to floating-point format.arm_dequantize_s16_f32functionDequantize an int16t array back to floating-point format.arm_strided_slice_s8functionStrided slice function for int8 data.arm_strided_slice_s16functionStrided slice function for int16 data.arm_strided_slice_s32functionStrided slice function for int32 data.arm_gather_s8functionGather elements along an axis for int8 tensors.arm_gather_s16functionGather elements along an axis for int16 tensors.arm_gather_nd_s8functionGathernd slices for int8 tensors.arm_gather_nd_s16functionGathernd slices for int16 tensors.arm_tile_s8functionTile an int8 tensor along each dimension.arm_tile_s16functionTile an int16 tensor along each dimension.arm_broadcast_to_s8functionBroadcast an int8 tensor to a target shape.arm_broadcast_to_s16functionBroadcast an int16 tensor to a target shape.arm_scatter_nd_s8functionScatter updates into a zero-initialized output tensor for int8.arm_scatter_nd_s16functionScatter updates into a zero-initialized output tensor for int16.arm_mirror_pad_s8functionMirror-pad an int8 tensor (REFLECT or SYMMETRIC mode).arm_mirror_pad_s16functionMirror-pad an int16 tensor (REFLECT or SYMMETRIC mode).arm_where_s8functionWHERE operator: return coordinates of non-zero elements in condition.arm_where_s16functionWHERE operator: return coordinates of non-zero elements in condition (int16).arm_select_v2_s8functionSELECTV2 with broadcast for int8 tensors.arm_select_v2_s16functionSELECTV2 with broadcast for int16 tensors.arm_reverse_sequence_s8functionReverse variable-length sequences along a dimension for int8.arm_reverse_sequence_s16functionReverse variable-length sequences along a dimension for int16.arm_dynamic_update_slice_s8functionUpdate a slice of an int8 operand tensor at runtime-determined indices.arm_dynamic_update_slice_s16functionUpdate a slice of an int16 operand tensor at runtime-determined indices.
s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
arm_cmsis_nn_status arm_convolve_wrapper_s4( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_wrapper_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 packed with 2x int4 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Get the required buffer size for armconvolvewrappers4.
int32_t arm_convolve_wrapper_s4_get_buffer_size( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required buffer size for arm_convolve_wrapper_s4.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Get the required buffer size for armconvolvewrappers4 for Arm(R) Helium Architecture case.
int32_t arm_convolve_wrapper_s4_get_buffer_size_mve( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required buffer size for arm_convolve_wrapper_s4 for Arm(R) Helium Architecture case.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Get the required buffer size for armconvolvewrappers4 for processors with DSP extension.
int32_t arm_convolve_wrapper_s4_get_buffer_size_dsp( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required buffer size for arm_convolve_wrapper_s4 for processors with DSP extension.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
arm_cmsis_nn_status arm_convolve_wrapper_s8( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
- On builds with ARM_MATH_MVEI (without ARM_MATH_AUTOVECTORIZE), a layer that would run
arm_convolve_s8()and is in the gate ofarm_convolve_s8_small_cin()orarm_convolve_s8_3x3_c16_s1()runs that entry instead, with the same result, scratch and weight sums. The input depth is checked first, so other layers skip both gates.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_wrapper_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. The selected kernel only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build: this wrapper dispatches to `arm_convolve_s8()`, `arm_convolve_1x1_s8()`, `arm_convolve_1x1_s8_fast()`, `arm_convolve_1_x_n_s8()` and `arm_convolve_1x1_out_s8()`. The buffer contents are consumed only on builds with the MVE extension (ARM_MATH_MVEI), and on those builds every one of those kernels diagnoses a NULL buf with ARM_CMSIS_NN_ARG_ERROR; on other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed that way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Get the required buffer size for armconvolvewrappers8.
int32_t arm_convolve_wrapper_s8_get_buffer_size( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required buffer size for arm_convolve_wrapper_s8.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Get the required buffer size for armconvolves8 for Arm(R) Helium Architecture case.
int32_t arm_convolve_s8_get_buffer_size_mve(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)Get the required buffer size for arm_convolve_s8 for Arm(R) Helium Architecture case.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas, so a shape whose dimension product overflows is only reported by the leg that actually computes a buffer for it.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Get the required buffer size for armconvolvewrappers8 for Arm(R) Helium Architecture case.
int32_t arm_convolve_wrapper_s8_get_buffer_size_mve( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required buffer size for arm_convolve_wrapper_s8 for Arm(R) Helium Architecture case.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
Get the required buffer size for armconvolvewrappers8 for processors with DSP extension.
int32_t arm_convolve_wrapper_s8_get_buffer_size_dsp( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required buffer size for arm_convolve_wrapper_s8 for processors with DSP extension.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid. |
s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
arm_cmsis_nn_status arm_convolve_wrapper_s16( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const cmsis_nn_bias_data *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_wrapper_s8_get_buffer_size will return the buffer_size if required The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const cmsis_nn_bias_data * | in | Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct. |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
s16 grouped convolution optimized for the case where filterdims->c == 1 and inputch == outputch (channel multiplier = 1).
arm_cmsis_nn_status arm_convolve_s16_group_ch_mult_1( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const cmsis_nn_bias_data *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)s16 grouped convolution optimized for the case where filter_dims->c == 1 and input_ch == output_ch (channel multiplier = 1).
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Function context (unused, pass NULL-initialised). |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (strides, dilations, pads, activation). |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info (multiplier and shift). |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int16_t * | in | Input data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, 1] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions (unused, may be zero-initialised). |
bias_data | const cmsis_nn_bias_data * | in | Optional bias struct (int32 or int64). May be NULL. |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success. |
Get the required buffer size for armconvolvewrappers16.
int32_t arm_convolve_wrapper_s16_get_buffer_size( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required buffer size for arm_convolve_wrapper_s16.
An out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Get the required buffer size for armconvolvewrappers16 for for processors with DSP extension.
int32_t arm_convolve_wrapper_s16_get_buffer_size_dsp( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required buffer size for arm_convolve_wrapper_s16 for for processors with DSP extension.
An out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Get the required buffer size for armconvolvewrappers16 for Arm(R) Helium Architecture case.
int32_t arm_convolve_wrapper_s16_get_buffer_size_mve( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required buffer size for arm_convolve_wrapper_s16 for Arm(R) Helium Architecture case.
An out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any ARM_NN_MAX() or sum, so it can never collapse into a plausible positive size.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Basic s4 convolution function.
arm_cmsis_nn_status arm_convolve_s4( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Basic s4 convolution function.
- Supported framework: TensorFlow Lite micro
- Additional memory is required for optimization. Refer to argument ‘ctx’ for details.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
filter_data | const int8_t * | in | Packed Filter data pointer. Data type: int8 packed with 2x int4 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Basic s4 convolution function with a requirement of even number of kernels.
arm_cmsis_nn_status arm_convolve_even_s4( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Basic s4 convolution function with a requirement of even number of kernels.
- Supported framework: TensorFlow Lite micro
- Additional memory is required for optimization. Refer to argument ‘ctx’ for details.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_even_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions. Note the product must be even. |
filter_data | const int8_t * | in | Packed Filter data pointer. Data type: int8 packed with 2x int4 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` if not for MVE |
Basic s8 convolution function.
arm_cmsis_nn_status arm_convolve_s8( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *upscale_dims, const cmsis_nn_dims *output_dims, int8_t *output_data)Basic s8 convolution function.
- Supported framework: TensorFlow Lite micro
- Additional memory is required for optimization. Refer to argument ‘ctx’ for details.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. For grouped convolution the entries run over all output_dims->c channels, groups laid out consecutively. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI); an unfilled buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, and a NULL buf is currently reported as ARM_CMSIS_NN_ARG_ERROR. On other builds this function currently derives the same quantity itself and does not read the context. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] where HK, WK and CK are the spatial filter dimensions. CK != C_IN is used for grouped convolution, in which case the required conditions are C_IN = N * CK and C_OUT = N * M for N groups of size M. |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions for transpose. Format: [H_UP, W_UP] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` |
s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model.
arm_cmsis_nn_status arm_convolve_s8_small_cin( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *upscale_dims, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model. It copies each kernel row with one predicated vector load and multiplies four output channels per step.
- The output is identical to
arm_convolve_s8(). The bias is read through the weight sums, whicharm_convolve_weight_sum()fills as forarm_convolve_s8(); bias_dims and bias_data are unused. - Gate: upscale_dims NULL, C_IN from 1 to 3 with CK equal to C_IN (one group), dilation 1 in both dimensions, WK and HK at least 1 with WK x C_IN at most 16 and HK x WK x C_IN at most 48, and C_OUT a positive multiple of 4. Stride, padding and batch count are as for
arm_convolve_s8(). - Scratch: ctx->buf holds
arm_convolve_s8_get_buffer_size()bytes (4 x 16 x ceil(HK x WK x C_IN / 16) on ARM_MATH_MVEI builds), the same asarm_convolve_s8(), and needs no alignment. - It is a direct entry:
arm_convolve_s8()does not call it, andarm_convolve_wrapper_s8()calls it for layers in the gate that it would otherwise pass toarm_convolve_s8(). A caller that selects the kernel per layer ahead of time calls it for layers in the gate andarm_convolve_s8()for every other layer, or onARM_CMSIS_NN_NO_IMPL_ERROR. Both take the same arguments, scratch and weight sums.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context with `arm_convolve_s8_get_buffer_size()` bytes of scratch, all of which may be written |
weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, as for `arm_convolve_s8()` |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters, as for `arm_convolve_s8()` |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions for transpose. Format: [H_UP, W_UP] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - an argument error that `arm_convolve_s8()` reports: ctx->buf is NULL, C_IN or C_OUT is not a multiple of the group count C_IN / CK, or weight_sum_ctx->buf is NULL on builds with ARM_MATH_MVEI. These are checked before the gate. `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer is outside the gate below, or the build lacks ARM_MATH_MVEI or defines ARM_MATH_AUTOVECTORIZE; nothing is written, to the output or to the scratch `ARM_CMSIS_NN_SUCCESS` - Successful operation |
s8 3x3 convolution over 16 input channels with unit stride.
arm_cmsis_nn_status arm_convolve_s8_3x3_c16_s1( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *upscale_dims, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 3x3 convolution over 16 input channels with unit stride. It reads the kernel rows of a patch inside the input in place, copying only patches that cross the border, and multiplies four output pixels per filter load.
- The output is identical to
arm_convolve_s8(). The bias is read through the weight sums; bias_dims and bias_data are unused. - Gate: upscale_dims NULL, C_IN and CK both 16 (one group), HK and WK both 3, and stride and dilation 1 in both dimensions. Padding, batch count and C_OUT are as for
arm_convolve_s8(). - It is a direct entry:
arm_convolve_s8()does not call it, andarm_convolve_wrapper_s8()calls it for layers in the gate that it would otherwise pass toarm_convolve_s8(). A caller that selects the kernel per layer ahead of time calls it for layers in the gate andarm_convolve_s8()for every other layer, or onARM_CMSIS_NN_NO_IMPL_ERROR. The gate does not overlap that ofarm_convolve_s8_small_cin().
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context with `arm_convolve_s8_get_buffer_size()` bytes of scratch (576 bytes on ARM_MATH_MVEI builds), all of which may be written |
weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, as for `arm_convolve_s8()` |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters, as for `arm_convolve_s8()` |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
upscale_dims | const cmsis_nn_dims * | in | Upscale tensor dimensions for transpose. Format: [H_UP, W_UP] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_convolve_s8_small_cin()` `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer is outside the gate below, or the build lacks ARM_MATH_MVEI or defines ARM_MATH_AUTOVECTORIZE; nothing is written, to the output or to the scratch `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Get the required buffer size for s4 convolution function.
int32_t arm_convolve_s4_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)Get the required buffer size for s4 convolution function.
The dimensions and the byte count are both checked here, so an out-of-range shape returns -1 on every build target rather than a wrapped size.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Get the required buffer size for armconvolveevens4.
int32_t arm_convolve_even_s4_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)Get the required buffer size for arm_convolve_even_s4.
Forwards to arm_convolve_s4_get_buffer_size(): the even_s4 kernel stages up to four im2col rows of filter_dims->w * filter_dims->h * input_dims->c int8 elements, byte-for-byte the size that sizer returns. The equality, including the -1 answers for out-of-range shapes, is pinned by a Unity test.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Get the required buffer size for s8 convolution function.
int32_t arm_convolve_s8_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)Get the required buffer size for s8 convolution function.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas, so a shape whose dimension product overflows is only reported by the leg that actually computes a buffer for it.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Get the required buffer size for s8 convolution and depthwise convolution weight sum.
int32_t arm_convolve_s8_get_weights_sum_size(const cmsis_nn_dims *output_dims)Get the required buffer size for s8 convolution and depthwise convolution weight sum.
For a valid (non-negative, in-range) output_dims->c, returns output_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. A negative or out-of-range output_dims->c returns -1 on builds with the MVE extension; elsewhere no weight sum buffer is used and the answer stays 0.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
output_dims | const cmsis_nn_dims * | in | Output (activation) tensor dimensions. Format: [N, H, W, C_COUT] |
Returns
| Description |
|---|
| The function returns required weight sum buffer size in bytes, or -1 if output_dims->c is negative or the required size would not fit in an int32_t |
Wrapper to select optimal transposed convolution algorithm depending on parameters.
arm_cmsis_nn_status arm_transpose_conv_wrapper_s8( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_context *reverse_conv_ctx, const cmsis_nn_transpose_conv_params *transpose_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Wrapper to select optimal transposed convolution algorithm depending on parameters.
- Supported framework: TensorFlow Lite micro
- Additional memory is required for optimization. Refer to arguments ‘ctx’ and ‘reverse_conv_ctx’ for details.
- Dilation is not supported: transpose_conv_params->dilation must be 1 in both dimensions, otherwise
ARM_CMSIS_NN_ARG_ERRORis returned.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. `arm_transpose_conv_s8_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. The function only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and transpose_conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing transpose_conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Compute the sums over filter_data exactly as passed to this function: this wrapper guarantees that whatever filter preparation it performs internally preserves the per-output-channel sums, so no reversed or otherwise rearranged copy of the weights is needed for this step. Pass a valid context on every build. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the reverse-convolution route forwards this context to `arm_convolve_s8()`; an unfilled buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, and a NULL buf is currently reported as ARM_CMSIS_NN_ARG_ERROR. On other builds the contents are currently not read. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
reverse_conv_ctx | const cmsis_nn_context * | in, out | Function context for the reversed filter used when this wrapper routes to the reverse convolution. Holds filter height * filter width * input channels * output channels int8 values; `arm_transpose_conv_s8_get_reverse_conv_buffer_size()` returns the required size (0 when the reverse-convolution route is not taken). The caller is expected to clear the buffer, if applicable, for security reasons. |
transpose_conv_params | const cmsis_nn_transpose_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of transpose_conv_params->input_offset : [-127, 128] Range of transpose_conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each out channel. |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Basic s8 transpose convolution function.
arm_cmsis_nn_status arm_transpose_conv_s8( const cmsis_nn_context *ctx, const cmsis_nn_context *output_ctx, const cmsis_nn_transpose_conv_params *transpose_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Basic s8 transpose convolution function.
- Supported framework: TensorFlow Lite micro
- Additional memory is required for optimization. Refer to argument ‘ctx’ for details; ‘output_ctx’ is unused.
- Dilation is not supported: transpose_conv_params->dilation must be 1 in both dimensions, otherwise
ARM_CMSIS_NN_ARG_ERRORis returned.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_transpose_conv_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
output_ctx | const cmsis_nn_context * | in, out | Not accessed by this function: its buffer is neither read nor written, and it therefore has no size requirement. The parameter exists only to keep one signature across the transpose-conv family, whose float twins ignore it the same way; `arm_transpose_conv_wrapper_s8()` forwards its reverse_conv_ctx into this slot. In-tree callers pass a valid context, whose buf may be NULL. |
transpose_conv_params | const cmsis_nn_transpose_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of transpose_conv_params->input_offset : [-127, 128] Range of transpose_conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each out channel. |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Get the required buffer size for ctx in s8 transpose conv function.
int32_t arm_transpose_conv_s8_get_buffer_size( const cmsis_nn_transpose_conv_params *transposed_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *out_dims)Get the required buffer size for ctx in s8 transpose conv function.
The returned size is safe for both arm_transpose_conv_s8() and arm_transpose_conv_wrapper_s8(): it is the larger of the two routes’ requirements, so it may exceed what the wrapper’s reverse-convolution route alone would need. When either route is out of range the sentinel is propagated ahead of that comparison, so -1 is never collapsed into a plausible positive size by the other route.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
transposed_conv_params | const cmsis_nn_transpose_conv_params * | in | Transposed convolution parameters |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
out_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative, either stride is not positive, or the required size would not fit in an int32_t |
Get the required buffer size for outputctx in s8 transpose conv function.
int32_t arm_transpose_conv_s8_get_reverse_conv_buffer_size( const cmsis_nn_transpose_conv_params *transposed_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)Get the required buffer size for output_ctx in s8 transpose conv function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
transposed_conv_params | const cmsis_nn_transpose_conv_params * | in | Transposed convolution parameters |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Get size of additional buffer required by armtransposeconvs8() for Arm(R) Helium Architecture case.
int32_t arm_transpose_conv_s8_get_buffer_size_mve( const cmsis_nn_transpose_conv_params *transposed_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *out_dims)Get size of additional buffer required by arm_transpose_conv_s8() for Arm(R) Helium Architecture case.
The returned size is safe for both arm_transpose_conv_s8() and arm_transpose_conv_wrapper_s8(): it is the larger of the two routes’ requirements, so it may exceed what the wrapper’s reverse-convolution route alone would need. When either route is out of range the sentinel is propagated ahead of that comparison, so -1 is never collapsed into a plausible positive size by the other route.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
transposed_conv_params | const cmsis_nn_transpose_conv_params * | in | Transposed convolution parameters |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
out_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative, either stride is not positive, or the required size would not fit in an int32_t |
Basic s16 convolution function.
arm_cmsis_nn_status arm_convolve_s16( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const cmsis_nn_bias_data *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)Basic s16 convolution function.
- Supported framework: TensorFlow Lite micro
- Additional memory is required for optimization. Refer to argument ‘ctx’ for details.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const cmsis_nn_bias_data * | in | Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct. |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` |
Pointwise s16 convolution function: no stride, no padding, no dilation.
arm_cmsis_nn_status arm_convolve_1x1_s16_ns_np_nd( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const cmsis_nn_bias_data *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)Pointwise s16 convolution function: no stride, no padding, no dilation.
- Supported framework: TensorFlow Lite micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const cmsis_nn_bias_data * | in | Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct. |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` |
armconvolves16fastsmallkernel function.
arm_cmsis_nn_status arm_convolve_s16_fast_small_kernel( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const cmsis_nn_bias_data *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)arm_convolve_s16_fast_small_kernel function. The kernel size is <=8
- Supported framework: TensorFlow Lite micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const cmsis_nn_bias_data * | in | Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct. |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` |
Get the required buffer size for s16 convolution function.
int32_t arm_convolve_s16_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)Get the required buffer size for s16 convolution function.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Fast s4 version for 1x1 convolution (non-square shape).
arm_cmsis_nn_status arm_convolve_1x1_s4_fast( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Fast s4 version for 1x1 convolution (non-square shape).
-
Supported framework : TensorFlow Lite Micro
-
The following constrains on the arguments apply
- conv_params->padding.w = conv_params->padding.h = 0
- conv_params->stride.w = conv_params->stride.h = 1
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_1x1_s4_fast_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 packed with 2x int4 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
s4 version for 1x1 convolution with support for non-unity stride values
arm_cmsis_nn_status arm_convolve_1x1_s4( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)s4 version for 1x1 convolution with support for non-unity stride values
-
Supported framework : TensorFlow Lite Micro
-
The following constrains on the arguments apply
- conv_params->padding.w = conv_params->padding.h = 0
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. None is required by this function. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 packed with 2x int4 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Fast s8 version for 1x1 convolution (non-square shape).
arm_cmsis_nn_status arm_convolve_1x1_s8_fast( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Fast s8 version for 1x1 convolution (non-square shape).
-
Supported framework : TensorFlow Lite Micro
-
The following constrains on the arguments apply
- conv_params->padding.w = conv_params->padding.h = 0
- conv_params->stride.w = conv_params->stride.h = 1
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_1x1_s8_fast_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. This function reads the buffer contents only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. On other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. Note also that on an Arm Compiler build (__ARMCC_VERSION >= 6010050) with ARM_MATH_DSP and without ARM_MATH_MVEI, supplying ctx->buf selects a buffered path that never reads weight_sum_ctx. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Get the required buffer size for armconvolve1x1s4fast.
int32_t arm_convolve_1x1_s4_fast_get_buffer_size(const cmsis_nn_dims *input_dims)Get the required buffer size for arm_convolve_1x1_s4_fast.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions |
Returns
| Description |
|---|
| The function returns the required buffer size in bytes, or -1 if input_dims->c is negative. No build needs this scratch buffer, so every valid shape returns 0. |
Get the required buffer size for armconvolve1x1s8fast.
int32_t arm_convolve_1x1_s8_fast_get_buffer_size(const cmsis_nn_dims *input_dims)Get the required buffer size for arm_convolve_1x1_s8_fast.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) dimensions |
Returns
| Description |
|---|
| The function returns the required buffer size in bytes, or -1 if input_dims->c is negative. On builds that need this scratch buffer it also returns -1 if the required size would not fit in an int32_t; other builds need no buffer and return 0. |
s8 version for 1x1 convolution with support for non-unity stride values
arm_cmsis_nn_status arm_convolve_1x1_s8( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 version for 1x1 convolution with support for non-unity stride values
-
Supported framework : TensorFlow Lite Micro
-
The following constrains on the arguments apply
- conv_params->padding.w = conv_params->padding.h = 0
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. None is required by this function. |
weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. This function reads the buffer contents only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. On other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
1xn convolution
arm_cmsis_nn_status arm_convolve_1_x_n_s8( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)1xn convolution
-
Supported framework : TensorFlow Lite Micro
-
The following constraints on the arguments apply
- input_dims->h, filter_dims->h and output_dims->h equal 1, and conv_params->padding.h is 0
- conv_params->dilation.w is 1 and conv_params->stride.w is positive
- conv_params->stride.w * input_dims->c is a multiple of 4
- conv_params->padding.w, input_dims->w and output_dims->w are not negative, and filter_dims->w is at least 1
-
Any horizontal padding and output width are handled, including an odd total padding, a filter wider than the input and a VALID layer whose stride leaves trailing input unused. On MVE builds the output columns whose window starts before or ends past the input read a padded copy of the input columns they span, staged in ctx; the other columns read the input in place.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_1_x_n_s8_get_buffer_size will return the buffer_size if required. buf must not be NULL. On builds with the MVE extension (ARM_MATH_MVEI) a non-zero ctx->size smaller than the staging the layer needs is rejected with ARM_CMSIS_NN_ARG_ERROR. The caller is expected to clear the buffer, if applicable, for security reasons. |
weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. The contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR; on other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Pre-computes per-output-channel weight sums for a standard convolution.
arm_cmsis_nn_status arm_convolve_weight_sum( int32_t *vector_sum_buf, const int8_t *rhs, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims, const int32_t lhs_offset, const int32_t *bias_data)Pre-computes per-output-channel weight sums for a standard convolution.
- Supported framework : TensorFlow Lite Micro
- The buffer pointed to by
vector_sum_bufmust be at leastoutput_dims->c × sizeof(int32_t)bytes.arm_convolve_s8_get_weights_sum_size()returns that size on builds that use the sums, 0 elsewhere, and -1 for an output_dims->c that is negative or too large to size. - Layout: one int32 per output channel, indexed 0..
output_dims->c - 1. Entry j holdslhs_offset * sum(weights of output channel j) + bias_data[j], i.e. the bias and the input-offset contribution folded together. For grouped convolution the entries run over all output channels, with the groups laid out consecutively. - This is the buffer the
weight_sum_ctxparameter of the s8 convolution kernels carries. Those kernels currently treat it as an input they only read, so it has to be filled before the call - see the individual functions for what each one currently does on MVE and non-MVE builds. - Reuse and invalidation: the contents depend only on
rhs,bias_dataandlhs_offset. They do not depend on the activations, so a buffer stays valid across calls and across batches for as long as those three are unchanged - for a static model the sums can be computed once at load time rather than per inference. Recompute whenever the weights, the bias or the input offset change (for example on requantization or a weight reload). The buffer is sized by one layer’soutput_dims->cand is specific to that layer’s weights, so it cannot be shared between layers; give each layer its own. - Returns
ARM_CMSIS_NN_NO_IMPL_ERRORon builds without the MVE extension, where the sums are currently not consumed.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
vector_sum_buf | int32_t * | out | Pointer to the buffer that will hold the weight sums. |
rhs | const int8_t * | in | Pointer to the filter weights. Data type: int8 |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, KH, KW, C_IN] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
lhs_offset | const int32_t | in | Input-offset added to every input element before MAC. Range: [-127, 128] |
bias_data | const int32_t * | in | Optional bias pointer. Data type: int32 |
Returns
| Description |
|---|
| `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments, `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed and the buffer is left untouched, or `ARM_CMSIS_NN_SUCCESS` on success. Portable code should not treat the `NO_IMPL_ERROR` case as a failure. |
Pre-computes per-channel weight sums for a depthwise convolution.
arm_cmsis_nn_status arm_depthwise_convolve_weight_sum( int32_t *vector_sum_buf, int8_t *scratch_buf, const int8_t *rhs, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims, const int32_t lhs_offset, const int32_t *bias_data)Pre-computes per-channel weight sums for a depthwise convolution.
- Supported framework : TensorFlow Lite Micro
- Layout: one int32 per channel, sized by
arm_convolve_s8_get_weights_sum_size(). Entry j holdsbias_data[j] + lhs_offset * sum(kernel values of channel j). - Reuse and invalidation follow the same rules as
arm_convolve_weight_sum(): the contents depend only onrhs,bias_dataandlhs_offset, so they may be computed once and reused until one of those changes, and they are specific to a single layer. - Returns
ARM_CMSIS_NN_NO_IMPL_ERRORon builds without the MVE extension, where the sums are currently not consumed. - Not interchangeable with
arm_convolve_weight_sum(): this function walks the channel-interleaved depthwise layout[1, KH, KW, C_OUT]with a stride of C_OUT, whereasarm_convolve_weight_sum()sums contiguous runs ofKH * KW * C_INweights. The two agree only by coincidence. Several in-tree tests do fill a depthwise weight_sum_ctx witharm_convolve_weight_sum()and are still correct, for one of three unrelated reasons:arm_depthwise_conv_wrapper_s8()does not consume the buffer on that route at all (ch_mult != 1, batches != 1, or a dilation the optimized route does not take - see that function); the wrapper converts the layer to a regular convolution, so conv-style sums are what is wanted; or C_OUT is 1, which collapses the stride-C_OUT walk to a contiguous one and makes the two helpers compute identical values. None of those generalise, so do not read them as licence to substitute one helper for the other. Use this function wherever the sums are actually read.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
vector_sum_buf | int32_t * | out | Buffer to hold the computed weight sums. |
scratch_buf | int8_t * | in, out | Currently unused: the implementation does not read or write it on any build, so NULL is accepted. Retained for signature compatibility; if a real buffer is passed, the caller is expected to clear it for security reasons. |
rhs | const int8_t * | in | Depthwise convolution weights. Data type: int8 |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise-convolution parameters (stride, dilation, pad, etc.) |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, KH, KW, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
lhs_offset | const int32_t | in | Input-offset applied before MAC. Range: [-127, 128] |
bias_data | const int32_t * | in | Optional bias pointer. Data type: int32 |
Returns
| Description |
|---|
| `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments, `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed and the buffer is left untouched, or `ARM_CMSIS_NN_SUCCESS` on success. Portable code should not treat the `NO_IMPL_ERROR` case as a failure. |
Optimised convolution for 1x1 output images (shape of BX1x1xCOUT) for 8x8 computations.
arm_cmsis_nn_status arm_convolve_1x1_out_s8( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Optimised convolution for 1x1 output images (shape of BX1x1xC_OUT) for 8x8 computations.
-
Supported framework : TensorFlow Lite Micro
-
Optimised for Bx1×1xC output CNN layers.
-
Constraints:
output_dims->handoutput_dims->wmust equal 1output_dims->cis expected to be a multiple of 4 for best performance
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that supplies a scratch buffer for activation rearrangement. A NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. The buffer must hold one 4-byte-aligned GEMM row, that is round_up_4(filter_dims->h * filter_dims->w * filter_dims->c) bytes, as returned by `arm_convolve_1x1_out_s8_get_buffer_size()`. The requirement does not scale with the group count: the kernel rewinds its im2col cursor to the start of the buffer after each group. Setting ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check, which is what TFLite Micro and derivatives do today. The caller is expected to clear the buffer, if applicable, for security reasons. |
weight_sum_ctx | const cmsis_nn_context * | in | Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. The contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR; on other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (stride, dilation, pad, offsets). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantisation multipliers and shifts. |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Pointer to input data. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, KH, KW, C_IN] |
filter_data | const int8_t * | in | Pointer to filter data. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, 1, 1, C_OUT] |
output_data | int8_t * | out | Pointer to output data. Data type: int8 |
Returns
| Description |
|---|
| `ARM_CMSIS_NN_ARG_ERROR` on bad args, or `ARM_CMSIS_NN_SUCCESS` on success. |
Get the required scratch buffer size for armconvolve1x1outs8().
int32_t arm_convolve_1x1_out_s8_get_buffer_size(const cmsis_nn_dims *filter_dims)Get the required scratch buffer size for arm_convolve_1x1_out_s8().
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, KH, KW, C_IN] |
Returns
| Description |
|---|
| For valid (non-negative, in-range) filter dimensions, the buffer size in bytes: round_up_4(KH * KW * C_IN) on builds with the MVE extension (ARM_MATH_MVEI), 0 otherwise, since `arm_convolve_1x1_out_s8()` only exists on MVE builds. Returns -1 if any of filter_dims->w, filter_dims->h or filter_dims->c is negative or out of int32_t range, or if the rounded-up product exceeds INT32_MAX. The validation runs on every build target, not just the MVE leg, so the contract does not vary by target. |
1xn convolution for s4 weights
arm_cmsis_nn_status arm_convolve_1_x_n_s4( const cmsis_nn_context *ctx, const cmsis_nn_conv_params *conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)1xn convolution for s4 weights
-
Supported framework : TensorFlow Lite Micro
-
The following constrains on the arguments apply
- stride.w * input_dims->c is a multiple of 4
- Explicit constraints(since it is for 1xN convolution) -## input_dims->h equals 1 -## output_dims->h equals 1 -## filter_dims->h equals 1
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. arm_convolve_1_x_n_s4_get_buffer_size will return the buffer_size if required The caller is expected to clear the buffer, if applicable, for security reasons. |
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 as packed int4 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Optional bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Get the required additional buffer size for 1xn convolution.
int32_t arm_convolve_1_x_n_s8_get_buffer_size( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required additional buffer size for 1xn convolution.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or conv_params->stride.w is not positive. On builds with the MVE extension (ARM_MATH_MVEI) that is the staging size of `arm_convolve_1_x_n_s8()`, at least filter W * C_IN bytes, or -1 if it would not fit in an int32_t; other builds return `arm_convolve_s8_get_buffer_size()`. |
Get the required additional buffer size for 1xn convolution.
int32_t arm_convolve_1_x_n_s4_get_buffer_size( const cmsis_nn_conv_params *conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get the required additional buffer size for 1xn convolution.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
conv_params | const cmsis_nn_conv_params * | in | Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or conv_params->stride.w is not positive. It also returns -1 if the required size would not fit in an int32_t; on a Helium build the route whose padding lines up with the stride needs no buffer and returns 0 without computing one. |
Wrapper function to pick the right optimized s8 depthwise convolution function.
arm_cmsis_nn_status arm_depthwise_conv_wrapper_s8( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Wrapper function to pick the right optimized s8 depthwise convolution function.
-
Supported framework: TensorFlow Lite
-
Picks one of the the following functions
arm_depthwise_conv_s8()arm_depthwise_conv_3x3_s8()- Cortex-M CPUs with DSP extension onlyarm_depthwise_conv_s8_opt()
-
Check details of
arm_depthwise_conv_s8_opt()for potential data that can be accessed outside of the boundary.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s8_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s8_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer, if applicable, for security reasons. |
weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, supplied by the caller. The selected kernel only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter, bias and dw_conv_params->input_offset are unchanged - see `arm_depthwise_convolve_weight_sum()` for the layout and the full reuse rules. Whether the buffer is consumed at all depends on the route this wrapper takes. It is forwarded to `arm_depthwise_conv_s8_opt()`, which reads it under MVE, only when dw_conv_params->ch_mult == 1, input_dims->n == 1, and either both dilations are 1 or the layer is 1D and dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, dw_conv_params->padding.h == 0, dilation.h == 1 and dilation.w >= 1. Such a dilated 1D layer therefore reads the sums too. Outside those cases the wrapper calls `arm_depthwise_conv_s8()`, which has no such parameter and ignores the context entirely - which is why several in-tree tests legitimately pass sums built by `arm_convolve_weight_sum()`, or none at all, on those routes (a 2D-dilated layer, for example). On MVE with input_dims->c == 1 and an output channel count above CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (8 on armclang, 1 otherwise), the layer is instead converted to a regular convolution, and conv-style sums from `arm_convolve_weight_sum()` are what that route wants. Where the sums are actually read, fill the buffer with `arm_depthwise_convolve_weight_sum()`, passing dw_conv_params->input_offset as lhs_offset and the same bias given here, so that entry j holds input_offset * sum(weights of channel j) + bias[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. On the `arm_depthwise_conv_s8_opt()` route, a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on builds where the buffer is actually read (ARM_MATH_DSP and ARM_MATH_MVEI both defined); on other builds the parameter is unread and NULL is accepted. On MVE with input_dims->c == 1 and an output channel count above CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (8 on armclang, 1 otherwise), this wrapper instead diverts to `arm_convolve_wrapper_s8()`. That diversion exists only on MVE, and every kernel it can dispatch to diagnoses a NULL buf with ARM_CMSIS_NN_ARG_ERROR, so that route is covered too. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions; see weight_sum_ctx for which dilated layers take the `arm_depthwise_conv_s8_opt()` route. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` on successful completion, or `ARM_CMSIS_NN_ARG_ERROR` on the `arm_depthwise_conv_s8_opt()` route if ctx->buf is NULL when a scratch buffer is required, or if weight_sum_ctx->buf is NULL on builds where it is read (ARM_MATH_DSP and ARM_MATH_MVEI both defined), or if ctx->size is non-zero and below `arm_depthwise_conv_s8_opt_get_buffer_size()` for a layer its channel path runs, or if that sizer returns -1 (a negative dimension or a byte count it cannot represent), or on the MVE `arm_convolve_wrapper_s8()` diversion route if weight_sum_ctx->buf is NULL. |
Wrapper function to pick the right optimized s4 depthwise convolution function.
arm_cmsis_nn_status arm_depthwise_conv_wrapper_s4( const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Wrapper function to pick the right optimized s4 depthwise convolution function.
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s4_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s4_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer ,if applicable, for security reasons. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43]. |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` - Successful completion. |
Get size of additional buffer required by armdepthwiseconvwrappers8().
int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get size of additional buffer required by arm_depthwise_conv_wrapper_s8().
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Get size of additional buffer required by armdepthwiseconvwrappers8() for processors with DSP extension.
int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get size of additional buffer required by arm_depthwise_conv_wrapper_s8() for processors with DSP extension.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Get size of additional buffer required by armdepthwiseconvwrappers8() for Arm(R) Helium Architecture case.
int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size_mve( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get size of additional buffer required by arm_depthwise_conv_wrapper_s8() for Arm(R) Helium Architecture case.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Get size of additional buffer required by armdepthwiseconvwrappers4().
int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get size of additional buffer required by arm_depthwise_conv_wrapper_s4().
This sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as arm_depthwise_conv_s8_opt_get_buffer_size(), so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Get size of additional buffer required by armdepthwiseconvwrappers4() for processors with DSP extension.
int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get size of additional buffer required by arm_depthwise_conv_wrapper_s4() for processors with DSP extension.
This sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as arm_depthwise_conv_s8_opt_get_buffer_size(), so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Get size of additional buffer required by armdepthwiseconvwrappers4() for Arm(R) Helium Architecture case.
int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size_mve( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get size of additional buffer required by arm_depthwise_conv_wrapper_s4() for Arm(R) Helium Architecture case.
This sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as arm_depthwise_conv_s8_opt_get_buffer_size(), so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Basic s8 depthwise convolution function that doesn't have any constraints on the input dimensions.
arm_cmsis_nn_status arm_depthwise_conv_s8( const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Basic s8 depthwise convolution function that doesn’t have any constraints on the input dimensions.
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s8_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s8()`, size the context with `arm_depthwise_conv_wrapper_s8_get_buffer_size()` instead, because another route through that wrapper does require a buffer. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Basic s4 depthwise convolution function that doesn't have any constraints on the input dimensions.
arm_cmsis_nn_status arm_depthwise_conv_s4( const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input, const cmsis_nn_dims *filter_dims, const int8_t *kernel, const cmsis_nn_dims *bias_dims, const int32_t *bias, const cmsis_nn_dims *output_dims, int8_t *output)Basic s4 depthwise convolution function that doesn’t have any constraints on the input dimensions.
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s4_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s4()`, size the context with `arm_depthwise_conv_wrapper_s4_get_buffer_size()` instead, because another route through that wrapper does require a buffer. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
input | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
kernel | const int8_t * | in | Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43]. |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output | int8_t * | in, out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Basic s16 depthwise convolution function that doesn't have any constraints on the input dimensions.
arm_cmsis_nn_status arm_depthwise_conv_s16( const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int64_t *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)Basic s16 depthwise convolution function that doesn’t have any constraints on the input dimensions.
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s16_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s16()`, size the context with `arm_depthwise_conv_wrapper_s16_get_buffer_size()` instead, because another route through that wrapper does require a buffer. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) conv_params->input_offset : Not used conv_params->output_offset : Not used |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Wrapper function to pick the right optimized s16 depthwise convolution function.
arm_cmsis_nn_status arm_depthwise_conv_wrapper_s16( const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int64_t *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)Wrapper function to pick the right optimized s16 depthwise convolution function.
-
Supported framework: TensorFlow Lite
-
Picks one of the the following functions
arm_depthwise_conv_s16()arm_depthwise_conv_fast_s16()- Cortex-M CPUs with DSP extension only
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s16_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s16_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer, if applicable, for security reasons. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions. When ch_mult == 1 and filter_dims->w * filter_dims->h < 512, `arm_depthwise_conv_fast_s16()` is used for an undilated layer and for a 1D layer dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, dw_conv_params->padding.h == 0, dilation.h == 1 and dilation.w >= 1. Other layers use `arm_depthwise_conv_s16()`. Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->output_offset : Not used |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` - Successful completion. |
Get size of additional buffer required by armdepthwiseconvwrappers16().
int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get size of additional buffer required by arm_depthwise_conv_wrapper_s16().
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Get size of additional buffer required by armdepthwiseconvwrappers16() for processors with DSP extension.
int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get size of additional buffer required by arm_depthwise_conv_wrapper_s16() for processors with DSP extension.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Get size of additional buffer required by armdepthwiseconvwrappers16() for Arm(R) Helium Architecture case.
int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size_mve( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Get size of additional buffer required by arm_depthwise_conv_wrapper_s16() for Arm(R) Helium Architecture case.
Where a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any ARM_NN_MAX() or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid. |
Optimized s16 depthwise convolution function with constraint that inchannel equals outchannel.
arm_cmsis_nn_status arm_depthwise_conv_fast_s16( const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int64_t *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)Optimized s16 depthwise convolution function with constraint that in_channel equals out_channel.
ARM_CMSIS_NN_SUCCESS - Successful operation
-
Supported framework: TensorFlow Lite
-
The following constraints on the arguments apply
- ch_mult == 1: the number of input channels equals the number of output channels
- filter_dims->w * filter_dims->h < MAX_COL_COUNT (512)
- dw_conv_params->dilation.h == 1 and dw_conv_params->dilation.w >= 1
-
Recommended when number of channels is 4 or greater.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. `arm_depthwise_conv_fast_s16_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation.w is honoured; dw_conv_params->dilation.h must be 1. dw_conv_params->input_offset : Not used dw_conv_params->output_offset : Not used |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - ctx-buff == NULL and `arm_depthwise_conv_fast_s16_get_buffer_size()` != 0 or input channel != output channel or filter_dims->w * filter_dims->h >= MAX_COL_COUNT (512) or dw_conv_params->dilation.h != 1 or dw_conv_params->dilation.w < 1 |
Get the required buffer size for optimized s16 depthwise convolution function with constraint that inchannel equals outchannel.
int32_t arm_depthwise_conv_fast_s16_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)Get the required buffer size for optimized s16 depthwise convolution function with constraint that in_channel equals out_channel.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below).
arm_cmsis_nn_status arm_depthwise_conv_3x3_s8( const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below).
-
Supported framework : TensorFlow Lite Micro
-
The following constrains on the arguments apply
- Number of input channel equals number of output channels
- Filter height and width equals 3
- Padding along x is either 0 or 1.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - Unsupported dimension of tensors - Unsupported pad size along the x axis `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Optimized s8 depthwise convolution function with constraint that inchannel equals outchannel.
arm_cmsis_nn_status arm_depthwise_conv_s8_opt( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Optimized s8 depthwise convolution function with constraint that in_channel equals out_channel.
-
Supported framework: TensorFlow Lite
-
The following constrains on the arguments apply
- Number of input channel equals number of output channels or ch_mult equals 1
-
Reccomended when number of channels is 4 or greater.
-
On builds with ARM_MATH_DSP and ARM_MATH_MVEI, layers that
arm_depthwise_conv_s8_opt_planar_supported()accepts run the planar path, with the same result asarm_depthwise_conv_s8_opt_planar(), unless ctx->size cannot hold its plane; every other layer runs the channel path ofarm_depthwise_conv_s8_opt_channelwise(). Callers that choose the path ahead of time can call either one directly.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer if required by the function. `arm_depthwise_conv_s8_opt_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons. |
weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, supplied by the caller and only read by this function. See the note below for how to size, fill and reuse the buffer and for when a NULL buf is diagnosed. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation.w is honoured; dw_conv_params->dilation.h must be 1. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - input channel != output channel or ch_mult != 1, or dw_conv_params->dilation.h != 1 or dw_conv_params->dilation.w < 1, or ctx->buf is NULL when a scratch buffer is required, or ctx->size is non-zero and below `arm_depthwise_conv_s8_opt_get_buffer_size()` for a layer the channel path runs, or that sizer returns -1 (a negative dimension or a byte count it cannot represent) on the channel path, or weight_sum_ctx->buf is NULL on builds where it is read (ARM_MATH_DSP and ARM_MATH_MVEI both defined) `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Whether armdepthwiseconvs8opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels.
int32_t arm_depthwise_conv_s8_opt_planar_supported( const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims)Whether arm_depthwise_conv_s8_opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels.
- The rule is plain C and evaluates the same on every build, so a code generator can apply it ahead of time. The planar path itself exists only on builds with ARM_MATH_DSP and ARM_MATH_MVEI.
- It depends only on the shapes and dw_conv_params: batch 1, C_IN equal to C_OUT, positive dimensions, stride 1, ch_mult 1, dilation.h 1, dilation.w at least 1 and at most 128 / C (integer division), at most 32 channels, the widths the path is faster for, and a plane that fits the scratch. This function is the reference for the rule; a mirror should be checked against it.
- The width thresholds follow measured speed and may be retuned in a later release. A caller that calls
arm_depthwise_conv_s8_opt_planar()directly must handleARM_CMSIS_NN_NO_IMPL_ERROR, for example by callingarm_depthwise_conv_s8_opt_channelwise().
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
Returns
| Description |
|---|
| 1 when the planar path takes the layer with a scratch of `arm_depthwise_conv_s8_opt_get_buffer_size_mve()` bytes, 0 otherwise. |
The planar path of armdepthwiseconvs8opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane.
arm_cmsis_nn_status arm_depthwise_conv_s8_opt_planar( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)The planar path of arm_depthwise_conv_s8_opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane.
- The output is identical to
arm_depthwise_conv_s8_opt()andarm_depthwise_conv_s8(). The bias is read through the weight sums; bias_dims and bias_data are unused.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context with the `arm_depthwise_conv_s8_opt_get_buffer_size()` scratch |
weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()` |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()` |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_depthwise_conv_s8_opt()`, except its channel-path ctx->size check: a ctx->size too small for the plane returns ARM_CMSIS_NN_NO_IMPL_ERROR instead `ARM_CMSIS_NN_NO_IMPL_ERROR` - `arm_depthwise_conv_s8_opt_planar_supported()` rejects the layer, ctx->size cannot hold its plane, or the build lacks ARM_MATH_DSP or ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation |
The channel-vectorized path of armdepthwiseconvs8opt() on its own, without the planar attempt.
arm_cmsis_nn_status arm_depthwise_conv_s8_opt_channelwise( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)The channel-vectorized path of arm_depthwise_conv_s8_opt() on its own, without the planar attempt.
- The output is identical to
arm_depthwise_conv_s8_opt()andarm_depthwise_conv_s8()for every layer.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context with the `arm_depthwise_conv_s8_opt_get_buffer_size()` scratch |
weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()` |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()` |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_depthwise_conv_s8_opt()` `ARM_CMSIS_NN_SUCCESS` - Successful operation |
s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of armdepthwiseconvs8opt(), for layers in its gate.
arm_cmsis_nn_status arm_depthwise_conv_s8_opt_3x3( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of arm_depthwise_conv_s8_opt(), for layers in its gate. It computes three output rows per weight load.
- The output is identical to
arm_depthwise_conv_s8_opt()andarm_depthwise_conv_s8(). The bias is read through the weight sums, whicharm_depthwise_convolve_weight_sum()fills as forarm_depthwise_conv_s8_opt(); bias_dims and bias_data are unused. - Gate: filter 3x3, dilation 1, N 1, C_IN equal to C_OUT with 16 <= C <= 2048 and C % 4 == 0, stride 1 or 2 and padding 0 or 1 in each dimension, input W >= 3 and H >= 1, output H >= 3 and W x H >= 16, every dimension at most 4096 and each tensor at most INT32_MAX elements, and the centre of the last output column’s window inside the input: (output W - 1) x stride.w - padding.w + 1 < input W. Input rows above or below the input count as padding, as in
arm_depthwise_conv_s8(). - Buffers: ctx and weight_sum_ctx and their buf are non-NULL, and ctx->size is at least
arm_depthwise_conv_s8_opt_3x3_get_buffer_size()(3008 + input W x C + 16 bytes). ctx->buf needs no alignment. - It is a direct entry:
arm_depthwise_conv_s8_opt()does not call it. A caller that selects the kernel per layer ahead of time callsarm_depthwise_conv_s8_opt_3x3_c64_s1()for C 64 with stride.h 1, this function for the rest of the gate, andarm_depthwise_conv_s8_opt()for every other layer, or onARM_CMSIS_NN_NO_IMPL_ERROR. All three take the same arguments and the same weight sums; a ctx that serves all three holds the larger ofarm_depthwise_conv_s8_opt_3x3_get_buffer_size()andarm_depthwise_conv_s8_opt_get_buffer_size(). On ARM_MATH_MVEI builds the second is the larger for a 3x3 filter when input W x C <= 1440.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context with at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` bytes of scratch |
weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()` |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()` |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer or a buffer is outside the gate below, or the build lacks ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation |
armdepthwiseconvs8opt3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps.
arm_cmsis_nn_status arm_depthwise_conv_s8_opt_3x3_c64_s1( const cmsis_nn_context *ctx, const cmsis_nn_context *weight_sum_ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)arm_depthwise_conv_s8_opt_3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps. It is the faster entry for those layers.
- The output is identical to
arm_depthwise_conv_s8_opt_3x3(),arm_depthwise_conv_s8_opt()andarm_depthwise_conv_s8(). The bias is read through the weight sums; bias_dims and bias_data are unused. - The two entries share only their gate, parameter packing and the code for the output_y % 3 remainder rows, so a build with -ffunction-sections and section garbage collection keeps only the code of the entries it calls.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context with at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` bytes of scratch |
weight_sum_ctx | const cmsis_nn_context * | in | Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()` |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()` |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_NO_IMPL_ERROR` - C_IN is not 64, dw_conv_params->stride.h is not 1, the layer or a buffer is outside the gate of `arm_depthwise_conv_s8_opt_3x3()`, or the build lacks ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Get the scratch size in bytes of armdepthwiseconvs8opt3x3() and armdepthwiseconvs8opt3x3c64s1().
int32_t arm_depthwise_conv_s8_opt_3x3_get_buffer_size(const cmsis_nn_dims *input_dims)Get the scratch size in bytes of arm_depthwise_conv_s8_opt_3x3() and arm_depthwise_conv_s8_opt_3x3_c64_s1().
- The size is the minimum ctx->size both entries accept. It depends only on the input width and channel count, since the filter is always 3x3.
- The function is plain C and returns the same size on every build, including builds without ARM_MATH_MVEI where the entries return
ARM_CMSIS_NN_NO_IMPL_ERROR, so a code generator can size the scratch ahead of time. A non-negative size is not a statement that the layer is in the gate of the entries.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN]. Only W and C_IN are read. |
Returns
| Description |
|---|
| 3008 + W x C_IN + 16 bytes, or -1 if W or C_IN is negative or the size would not fit in an int32_t |
Optimized s4 depthwise convolution function with constraint that inchannel equals outchannel.
arm_cmsis_nn_status arm_depthwise_conv_s4_opt( const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params *dw_conv_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Optimized s4 depthwise convolution function with constraint that in_channel equals out_channel.
-
Supported framework: TensorFlow Lite
-
The following constrains on the arguments apply
- Number of input channel equals number of output channels or ch_mult equals 1
-
Reccomended when number of channels is 4 or greater.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context that contains the additional buffer required by the function. `arm_depthwise_conv_s4_opt_get_buffer_size()` will return the buffer_size. A NULL ctx->buf is diagnosed with `ARM_CMSIS_NN_ARG_ERROR`. The caller is expected to clear the buffer, if applicable, for security reasons. |
dw_conv_params | const cmsis_nn_dw_conv_params * | in | Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used. |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43]. |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - input channel != output channel or ch_mult != 1 `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel.
int32_t arm_depthwise_conv_s8_opt_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel.
The dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t |
Get the required buffer size for optimized s4 depthwise convolution function with constraint that inchannel equals outchannel.
int32_t arm_depthwise_conv_s4_opt_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)Get the required buffer size for optimized s4 depthwise convolution function with constraint that in_channel equals out_channel.
The dimensions are not checked here: the query routes straight to the s8 _mve/_dsp leg and relies on the range checks inside that leg. Both legs apply the same check as arm_depthwise_conv_s8_opt_get_buffer_size(), so the answer for an out-of-range shape is the same on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used. |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [1, H, W, C_OUT] |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if input_dims->c or a filter dimension it reads is negative, or the required size would not fit in an int32_t. |
Basic s4 Fully Connected function.
arm_cmsis_nn_status arm_fully_connected_s4( const cmsis_nn_context *ctx, const cmsis_nn_fc_params *fc_params, const cmsis_nn_per_tensor_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Basic s4 Fully Connected function.
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_fully_connected_s4_get_buffer_size(). Do not size this context with `arm_fully_connected_s8_get_buffer_size()`: that sizes the kernel-sum buffer of a different kernel and does not describe this argument. |
fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_tensor_quant_params * | in | Per-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor. |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43]. |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Basic s8 Fully Connected function.
arm_cmsis_nn_status arm_fully_connected_s8( const cmsis_nn_context *ctx, const cmsis_nn_fc_params *fc_params, const cmsis_nn_per_tensor_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Basic s8 Fully Connected function.
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Per-output-channel kernel sums, supplied by the caller - not scratch memory that this function fills in. The library never populates ctx->buf for this function, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds this function adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. Sized by `arm_fully_connected_s8_get_buffer_size()`: filter_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call. |
fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_tensor_quant_params * | in | Per-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor. |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Basic s8 Fully Connected function using per channel quantization.
arm_cmsis_nn_status arm_fully_connected_per_channel_s8( const cmsis_nn_context *ctx, const cmsis_nn_fc_params *fc_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)Basic s8 Fully Connected function using per channel quantization.
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Per-output-channel kernel sums, supplied by the caller - not scratch memory that this function fills in. The library never populates ctx->buf for this function, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds this function adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. There is no per-channel sizer; `arm_fully_connected_s8_get_buffer_size()` returns the same quantity this function needs, filter_dims->c * sizeof(int32_t) where the sums are used and 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call. |
fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
s8 Fully Connected layer wrapper function
arm_cmsis_nn_status arm_fully_connected_wrapper_s8( const cmsis_nn_context *ctx, const cmsis_nn_fc_params *fc_params, const cmsis_nn_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 Fully Connected layer wrapper function
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Per-output-channel kernel sums, supplied by the caller - not scratch memory that this wrapper fills in. The library never populates ctx->buf here, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. The context is passed straight through to `arm_fully_connected_per_channel_s8()` or `arm_fully_connected_s8()` depending on quant_params->is_per_channel, and both read it the same way. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. Neither selected kernel writes the buffer, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds the selected kernel adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. There is no wrapper sizer; `arm_fully_connected_s8_get_buffer_size()` returns the quantity both routes need, filter_dims->c * sizeof(int32_t) where the sums are used and 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call. |
fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127] |
quant_params | const cmsis_nn_quant_params * | in | Per-channel or per-tensor quantization info. Check struct defintion for details. It contains the multiplier and shift value(s) to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
bias_data | const int32_t * | in | Bias data pointer. Data type: int32 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s32 biasdata.
arm_cmsis_nn_status arm_vector_sum_s8( int32_t *vector_sum_buf, const int32_t vector_cols, const int32_t vector_rows, const int8_t *vector_data, const int32_t lhs_offset, const int32_t rhs_offset, const int32_t *bias_data)Calculate the sum of each row in vector_data, multiply by lhs_offset and optionally add s32 bias_data.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
vector_sum_buf | int32_t * | in, out | Buffer for vector sums |
vector_cols | const int32_t | in | Number of vector columns |
vector_rows | const int32_t | in | Number of vector rows |
vector_data | const int8_t * | in | Vector of weigths data |
lhs_offset | const int32_t | in | Constant multiplied with each sum |
rhs_offset | const int32_t | in | Constant added to each vector element before sum |
bias_data | const int32_t * | in | Vector of bias data, added to each sum. |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s64 biasdata.
arm_cmsis_nn_status arm_vector_sum_s8_s64( int64_t *vector_sum_buf, const int32_t vector_cols, const int32_t vector_rows, const int8_t *vector_data, const int32_t lhs_offset, const int64_t *bias_data)Calculate the sum of each row in vector_data, multiply by lhs_offset and optionally add s64 bias_data.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
vector_sum_buf | int64_t * | in, out | Buffer for vector sums |
vector_cols | const int32_t | in | Number of vector columns |
vector_rows | const int32_t | in | Number of vector rows |
vector_data | const int8_t * | in | Vector of weigths data |
lhs_offset | const int32_t | in | Constant multiplied with each sum |
bias_data | const int64_t * | in | Vector of bias data, added to each sum. |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Get size of additional buffer required by armfullyconnecteds8().
int32_t arm_fully_connected_s8_get_buffer_size(const cmsis_nn_dims *filter_dims)Get size of additional buffer required by arm_fully_connected_s8(). See also arm_vector_sum_s8, which is required if buffer size is > 0.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | dimension of filter |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Get size of additional buffer required by armfullyconnecteds8() for processors with DSP extension.
int32_t arm_fully_connected_s8_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)Get size of additional buffer required by arm_fully_connected_s8() for processors with DSP extension.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | dimension of filter |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Get size of additional buffer required by armfullyconnecteds8() for Arm(R) Helium Architecture case.
int32_t arm_fully_connected_s8_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)Get size of additional buffer required by arm_fully_connected_s8() for Arm(R) Helium Architecture case.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | dimension of filter |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Basic s16 Fully Connected function.
arm_cmsis_nn_status arm_fully_connected_s16( const cmsis_nn_context *ctx, const cmsis_nn_fc_params *fc_params, const cmsis_nn_per_tensor_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int64_t *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)Basic s16 Fully Connected function.
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused. This function currently ignores the context entirely on every build - it neither reads nor writes ctx->buf - and `arm_fully_connected_s16_get_buffer_size()` returns 0 accordingly, so { NULL, 0 } is accepted. Unlike the s8 variants, no precomputed kernel sums are required here. None of this is a guarantee about future versions. |
fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. fc_params->input_offset : 0 fc_params->filter_offset : 0 fc_params->output_offset : 0 |
quant_params | const cmsis_nn_per_tensor_quant_params * | in | Per-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor. |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Basic s16 Fully Connected function using per channel quantization.
arm_cmsis_nn_status arm_fully_connected_per_channel_s16( const cmsis_nn_context *ctx, const cmsis_nn_fc_params *fc_params, const cmsis_nn_per_channel_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *kernel, const cmsis_nn_dims *bias_dims, const int64_t *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)Basic s16 Fully Connected function using per channel quantization.
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Scratch buffer that this function writes before it reads, on every build. It is filled here with one reduced int32 multiplier per output channel derived from quant_params->multiplier, so the caller supplies the storage only and the incoming contents are never used. Unlike the s8 variants, no precomputed kernel sums are expected, and clearing the buffer is harmless. Required on every build, not only under MVE: ctx or ctx->buf being NULL is reported as ARM_CMSIS_NN_ARG_ERROR, as is a non-zero ctx->size smaller than the requirement. A ctx->size of 0 is treated as undeclared and is not checked. Sized by `arm_fully_connected_per_channel_s16_get_buffer_size()`: filter_dims->c * sizeof(int32_t), which equals the output_dims->c entries written. The caller is expected to clear the buffer afterwards, if applicable, for security reasons. |
fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : 0 fc_params->filter_offset : 0 Range of fc_params->output_offset : 0 |
quant_params | const cmsis_nn_per_channel_quant_params * | in | Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
kernel | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
s16 Fully Connected layer wrapper function
arm_cmsis_nn_status arm_fully_connected_wrapper_s16( const cmsis_nn_context *ctx, const cmsis_nn_fc_params *fc_params, const cmsis_nn_quant_params *quant_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const int8_t *filter_data, const cmsis_nn_dims *bias_dims, const int64_t *bias_data, const cmsis_nn_dims *output_dims, int16_t *output_data)s16 Fully Connected layer wrapper function
- Supported framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Scratch buffer, whose use depends on the route taken. Unlike the s8 wrapper, no precomputed kernel sums are expected on either route, and clearing the buffer is harmless. When quant_params->is_per_channel is set, the context is passed to `arm_fully_connected_per_channel_s16()`, which writes it before reading it, on every build: it is filled there with one reduced int32 multiplier per output channel, so the caller supplies the storage only. On that route ctx or ctx->buf being NULL is reported as ARM_CMSIS_NN_ARG_ERROR, as is a non-zero ctx->size smaller than filter_dims->c * sizeof(int32_t); a ctx->size of 0 is treated as undeclared and is not checked. Size it with `arm_fully_connected_per_channel_s16_get_buffer_size()`. Otherwise the context goes to `arm_fully_connected_s16()`, which currently ignores it entirely, so { NULL, 0 } is accepted on that route. A caller that does not know the route in advance should size for the per-channel case, since `arm_fully_connected_s16_get_buffer_size()` returns 0. None of this is a guarantee about future versions. The caller is expected to clear the buffer afterwards, if applicable, for security reasons. |
fc_params | const cmsis_nn_fc_params * | in | Fully Connected layer parameters. Range of fc_params->input_offset : 0 fc_params->filter_offset : 0 Range of fc_params->output_offset : 0 |
quant_params | const cmsis_nn_quant_params * | in | Per-channel or per-tensor quantization info. Check struct defintion for details. It contains the multiplier and shift value(s) to be applied to each output channel |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN) |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used |
filter_data | const int8_t * | in | Filter data pointer. Data type: int8 |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used |
bias_data | const int64_t * | in | Bias data pointer. Data type: int64 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used. |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Get size of additional buffer required by armfullyconnecteds16().
int32_t arm_fully_connected_s16_get_buffer_size(const cmsis_nn_dims *filter_dims)Get size of additional buffer required by arm_fully_connected_s16().
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | dimension of filter |
Returns
| Description |
|---|
| The function returns required buffer size in bytes |
Get size of additional buffer required by armfullyconnecteds16() for processors with DSP extension.
int32_t arm_fully_connected_s16_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)Get size of additional buffer required by arm_fully_connected_s16() for processors with DSP extension.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | dimension of filter |
Returns
| Description |
|---|
| The function returns required buffer size in bytes |
Get size of additional buffer required by armfullyconnecteds16() for Arm(R) Helium Architecture case.
int32_t arm_fully_connected_s16_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)Get size of additional buffer required by arm_fully_connected_s16() for Arm(R) Helium Architecture case.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | dimension of filter |
Returns
| Description |
|---|
| The function returns required buffer size in bytes |
Get size of additional buffer required by armfullyconnectedperchannels16().
int32_t arm_fully_connected_per_channel_s16_get_buffer_size(const cmsis_nn_dims *filter_dims)Get size of additional buffer required by arm_fully_connected_per_channel_s16().
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | dimension of filter |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Get size of additional buffer required by armfullyconnectedperchannels16() for processors with DSP extension.
int32_t arm_fully_connected_per_channel_s16_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)Get size of additional buffer required by arm_fully_connected_per_channel_s16() for processors with DSP extension.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | dimension of filter |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
Get size of additional buffer required by armfullyconnectedperchannels16() for Arm(R) Helium Architecture case.
int32_t arm_fully_connected_per_channel_s16_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)Get size of additional buffer required by arm_fully_connected_per_channel_s16() for Arm(R) Helium Architecture case.
For a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
filter_dims | const cmsis_nn_dims * | in | dimension of filter |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t |
s8 elementwise add of two tensors with support for broadcasting.
arm_cmsis_nn_status arm_add_s8( const int8_t *input1_data, const cmsis_nn_dims *input1_dims, const int8_t *input2_data, const cmsis_nn_dims *input2_dims, const int32_t input1_offset, const int32_t input1_mult, const int32_t input1_shift, const int32_t input2_offset, const int32_t input2_mult, const int32_t input2_shift, const int32_t left_shift, int8_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max)s8 elementwise add of two tensors with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input1_data | const int8_t * | in | pointer to input tensor 1 |
input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
input2_data | const int8_t * | in | pointer to input tensor 2 |
input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input1_mult | const int32_t | in | multiplier for input 1 |
input1_shift | const int32_t | in | shift for input 1 |
input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input2_mult | const int32_t | in | multiplier for input 2 |
input2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output_data | int8_t * | out | pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
s8 elementwise add of scalar and vector
arm_cmsis_nn_status arm_add_scalar_s8( const int8_t *input_1_vect, const int8_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int8_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s8 elementwise add of scalar and vector
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int8_t * | in | pointer to input scalar |
input_2_vect | const int8_t * | in | pointer to input vector |
input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int8_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s8 elementwise add of two vectors
arm_cmsis_nn_status arm_elementwise_add_s8( const int8_t *input_1_vect, const int8_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int8_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s8 elementwise add of two vectors
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int8_t * | in | pointer to input vector 1 |
input_2_vect | const int8_t * | in | pointer to input vector 2 |
input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | input left shift. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int8_t * | in, out | pointer to output vector |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s8 elementwise absolute value
arm_cmsis_nn_status arm_abs_s8( const int8_t *input, const int32_t input_offset, int8_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const bool needs_rescale, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s8 elementwise absolute value
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | pointer to input vector |
input_offset | const int32_t | in | input offset |
output | int8_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
needs_rescale | const bool | in | indicates if output requantization is needed |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s8 elementwise square root
arm_cmsis_nn_status arm_sqrt_s8( const int8_t *input, const cmsis_nn_dims *input_dims, int8_t *output, const int8_t *sqrt_lut)s8 elementwise square root
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | pointer to input vector |
input_dims | const cmsis_nn_dims * | in | pointer to input tensor dimensions |
output | int8_t * | out | pointer to output vector |
sqrt_lut | const int8_t * | in | pointer to 256-entry lookup table |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise square root using piecewise LUT with linear interpolation
arm_cmsis_nn_status arm_sqrt_s16( const int16_t *input, const cmsis_nn_dims *input_dims, int16_t *output, const int16_t *sqrt_lut)s16 elementwise square root using piecewise LUT with linear interpolation
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | pointer to input vector |
input_dims | const cmsis_nn_dims * | in | pointer to input tensor dimensions |
output | int16_t * | out | pointer to output vector |
sqrt_lut | const int16_t * | in | pointer to 513-entry lookup table (int16_t) |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise square root without a lookup table
arm_cmsis_nn_status arm_sqrt_s16_tablefree( const int16_t *input, const cmsis_nn_dims *input_dims, int16_t *output, const float scale)s16 elementwise square root without a lookup table
Approximates output[i] = trunc(sqrt(input[i] * scale)) saturated to 32767, which is LiteRT’s int16 SQRT (dequantize in float32, sqrtf, divide by the output scale, truncate, clamp) for zero points 0, to within 1 LSB of LiteRT at every non-negative input for input scales 1e-7 to 1e-1 and output scales from 0.01x to 10x the full-range scale, saturating ones included. Inputs at or below 0 produce
- Needs no table; the int16 API does not depend on ARM_NN_ENABLE_F32/F16, and on targets without a floating-point unit the plain C path uses fmaf from the C library.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | pointer to input vector |
input_dims | const cmsis_nn_dims * | in | pointer to input tensor dimensions |
output | int16_t * | out | pointer to output vector |
scale | const float | in | input_scale / (output_scale * output_scale) as float32: take the float32-rounded tensor scales, evaluate in float64 and round once to float32. Must be finite and greater than 0. |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise absolute value
arm_cmsis_nn_status arm_abs_s16( const int16_t *input, const int32_t input_offset, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const bool needs_rescale, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s16 elementwise absolute value
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | pointer to input vector |
input_offset | const int32_t | in | input offset |
output | int16_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
needs_rescale | const bool | in | indicates if output requantization is needed |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
INT16 reciprocal square root using a per-operator LUT.
arm_cmsis_nn_status arm_rsqrt_s16_per_op( const int16_t *input, const int32_t input_offset, int16_t *output, const int32_t out_offset, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size, const int16_t *lut)INT16 reciprocal square root using a per-operator LUT.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input buffer. |
input_offset | const int32_t | in | Input tensor zero offset. The kernel evaluates each element as `input - input_offset` before the LUT lookup. |
output | int16_t * | out | Pointer to the output buffer. |
out_offset | const int32_t | in | Output tensor zero offset. |
out_activation_min | const int32_t | in | Minimum output clamp. |
out_activation_max | const int32_t | in | Maximum output clamp. |
block_size | const int32_t | in | Number of elements. |
lut | const int16_t * | in | Pointer to a 513-entry INT16 LUT in output domain. |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
INT16 reciprocal square root using a shared universal LUT.
arm_cmsis_nn_status arm_rsqrt_s16_universal( const int16_t *input, const int32_t input_offset, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const bool needs_rescale, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size, const int32_t *lut)INT16 reciprocal square root using a shared universal LUT.
In universal mode all RSQRT operators share a single LUT that captures the base 1/sqrt(x) shape, and operator-specific quantization is applied afterward via out_mult / out_shift. Because this two-step process introduces extra rounding stages, the output may differ from the per-op variant (arm_rsqrt_s16_per_op) by up to ±3 LSB per element. This is expected and acceptable for deployment.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input buffer. |
input_offset | const int32_t | in | Input tensor zero offset. The kernel evaluates each element as `input - input_offset` before the LUT lookup. |
output | int16_t * | out | Pointer to the output buffer. |
out_offset | const int32_t | in | Output tensor zero offset. |
out_mult | const int32_t | in | Output requantization multiplier. |
out_shift | const int32_t | in | Output requantization shift. |
needs_rescale | const bool | in | Whether requantization is required. |
out_activation_min | const int32_t | in | Minimum output clamp. |
out_activation_max | const int32_t | in | Maximum output clamp. |
block_size | const int32_t | in | Number of elements. |
lut | const int32_t * | in | Pointer to a 513-entry INT32 shared LUT in Q30 domain. |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR. |
s8 elementwise subtraction of two tensors with support for broadcasting.
arm_cmsis_nn_status arm_sub_s8( const int8_t *input1_data, const cmsis_nn_dims *input1_dims, const int8_t *input2_data, const cmsis_nn_dims *input2_dims, const int32_t input1_offset, const int32_t input1_mult, const int32_t input1_shift, const int32_t input2_offset, const int32_t input2_mult, const int32_t input2_shift, const int32_t left_shift, int8_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max)s8 elementwise subtraction of two tensors with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input1_data | const int8_t * | in | pointer to input tensor 1 |
input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
input2_data | const int8_t * | in | pointer to input tensor 2 |
input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input1_mult | const int32_t | in | multiplier for input 1 |
input1_shift | const int32_t | in | shift for input 1 |
input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input2_mult | const int32_t | in | multiplier for input 2 |
input2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output_data | int8_t * | out | pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
s8 elementwise subtract of scalar and vector (scalar - vector)
arm_cmsis_nn_status arm_sub_scalar_s8( const int8_t *input_1_vect, const int8_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int8_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s8 elementwise subtract of scalar and vector (scalar - vector)
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int8_t * | in | pointer to input scalar |
input_2_vect | const int8_t * | in | pointer to input vector |
input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int8_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s8 elementwise subtract of two vectors
arm_cmsis_nn_status arm_elementwise_sub_s8( const int8_t *input_1_vect, const int8_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int8_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s8 elementwise subtract of two vectors
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int8_t * | in | pointer to input vector 1 |
input_2_vect | const int8_t * | in | pointer to input vector 2 |
input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | input left shift. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int8_t * | in, out | pointer to output vector |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise add of two tensors with support for broadcasting.
arm_cmsis_nn_status arm_add_s16( const int16_t *input1_data, const cmsis_nn_dims *input1_dims, const int16_t *input2_data, const cmsis_nn_dims *input2_dims, const int32_t input1_offset, const int32_t input1_mult, const int32_t input1_shift, const int32_t input2_offset, const int32_t input2_mult, const int32_t input2_shift, const int32_t left_shift, int16_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max)s16 elementwise add of two tensors with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input1_data | const int16_t * | in | pointer to input tensor 1 |
input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
input2_data | const int16_t * | in | pointer to input tensor 2 |
input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input1_mult | const int32_t | in | multiplier for input 1 |
input1_shift | const int32_t | in | shift for input 1 |
input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input2_mult | const int32_t | in | multiplier for input 2 |
input2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output_data | int16_t * | out | pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
s16 elementwise add of scalar and vector
arm_cmsis_nn_status arm_add_scalar_s16( const int16_t *input_1_vect, const int16_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s16 elementwise add of scalar and vector
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int16_t * | in | pointer to input scalar |
input_2_vect | const int16_t * | in | pointer to input vector |
input_1_offset | const int32_t | in | offset for input 1. Not used. |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Not used. |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int16_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset. Not used. |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise add of two vectors
arm_cmsis_nn_status arm_elementwise_add_s16( const int16_t *input_1_vect, const int16_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s16 elementwise add of two vectors
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int16_t * | in | pointer to input vector 1 |
input_2_vect | const int16_t * | in | pointer to input vector 2 |
input_1_offset | const int32_t | in | offset for input 1. Not used. |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Not used. |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | input left shift. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int16_t * | in, out | pointer to output vector |
out_offset | const int32_t | in | output offset. Not used. |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise subtraction of two tensors with support for broadcasting.
arm_cmsis_nn_status arm_sub_s16( const int16_t *input1_data, const cmsis_nn_dims *input1_dims, const int16_t *input2_data, const cmsis_nn_dims *input2_dims, const int32_t input1_offset, const int32_t input1_mult, const int32_t input1_shift, const int32_t input2_offset, const int32_t input2_mult, const int32_t input2_shift, const int32_t left_shift, int16_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max)s16 elementwise subtraction of two tensors with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input1_data | const int16_t * | in | pointer to input tensor 1 |
input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
input2_data | const int16_t * | in | pointer to input tensor 2 |
input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input1_mult | const int32_t | in | multiplier for input 1 |
input1_shift | const int32_t | in | shift for input 1 |
input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input2_mult | const int32_t | in | multiplier for input 2 |
input2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output_data | int16_t * | out | pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
s16 elementwise subtract of scalar and vector (scalar - vector)
arm_cmsis_nn_status arm_sub_scalar_s16( const int16_t *input_1_vect, const int16_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s16 elementwise subtract of scalar and vector (scalar - vector)
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int16_t * | in | pointer to input scalar |
input_2_vect | const int16_t * | in | pointer to input vector |
input_1_offset | const int32_t | in | offset for input 1. Not used. |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Not used. |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int16_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset. Not used. |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise subtract of two vectors
arm_cmsis_nn_status arm_elementwise_sub_s16( const int16_t *input_1_vect, const int16_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s16 elementwise subtract of two vectors
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int16_t * | in | pointer to input vector 1 |
input_2_vect | const int16_t * | in | pointer to input vector 2 |
input_1_offset | const int32_t | in | offset for input 1. Not used. |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Not used. |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | input left shift. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int16_t * | in, out | pointer to output vector |
out_offset | const int32_t | in | output offset. Not used. |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s8 elementwise squared difference of two tensors with support for broadcasting.
arm_cmsis_nn_status arm_squared_difference_s8( const int8_t *input1_data, const cmsis_nn_dims *input1_dims, const int8_t *input2_data, const cmsis_nn_dims *input2_dims, const int32_t input1_offset, const int32_t input1_mult, const int32_t input1_shift, const int32_t input2_offset, const int32_t input2_mult, const int32_t input2_shift, const int32_t left_shift, int8_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max)s8 elementwise squared difference of two tensors with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input1_data | const int8_t * | in | pointer to input tensor 1 |
input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
input2_data | const int8_t * | in | pointer to input tensor 2 |
input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input1_mult | const int32_t | in | multiplier for input 1 |
input1_shift | const int32_t | in | shift for input 1 |
input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input2_mult | const int32_t | in | multiplier for input 2 |
input2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output_data | int8_t * | out | pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
s8 elementwise squared difference of scalar and vector.
arm_cmsis_nn_status arm_squared_difference_scalar_s8( const int8_t *input_1_vect, const int8_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int8_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s8 elementwise squared difference of scalar and vector.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int8_t * | in | pointer to input scalar |
input_2_vect | const int8_t * | in | pointer to input vector |
input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int8_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s8 elementwise squared difference of two vectors.
arm_cmsis_nn_status arm_elementwise_squared_difference_s8( const int8_t *input_1_vect, const int8_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int8_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s8 elementwise squared difference of two vectors.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int8_t * | in | pointer to input vector 1 |
input_2_vect | const int8_t * | in | pointer to input vector 2 |
input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int8_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise squared difference of two tensors with support for broadcasting.
arm_cmsis_nn_status arm_squared_difference_s16( const int16_t *input1_data, const cmsis_nn_dims *input1_dims, const int16_t *input2_data, const cmsis_nn_dims *input2_dims, const int32_t input1_offset, const int32_t input1_mult, const int32_t input1_shift, const int32_t input2_offset, const int32_t input2_mult, const int32_t input2_shift, const int32_t left_shift, int16_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max)s16 elementwise squared difference of two tensors with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input1_data | const int16_t * | in | pointer to input tensor 1 |
input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
input2_data | const int16_t * | in | pointer to input tensor 2 |
input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
input1_offset | const int32_t | in | offset for input 1 |
input1_mult | const int32_t | in | multiplier for input 1 |
input1_shift | const int32_t | in | shift for input 1 |
input2_offset | const int32_t | in | offset for input 2 |
input2_mult | const int32_t | in | multiplier for input 2 |
input2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output_data | int16_t * | out | pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
out_offset | const int32_t | in | output offset |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
s16 elementwise squared difference of scalar and vector.
arm_cmsis_nn_status arm_squared_difference_scalar_s16( const int16_t *input_1_vect, const int16_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s16 elementwise squared difference of scalar and vector.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int16_t * | in | pointer to input scalar |
input_2_vect | const int16_t * | in | pointer to input vector |
input_1_offset | const int32_t | in | offset for input 1 |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2 |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int16_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise squared difference of two vectors.
arm_cmsis_nn_status arm_elementwise_squared_difference_s16( const int16_t *input_1_vect, const int16_t *input_2_vect, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s16 elementwise squared difference of two vectors.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int16_t * | in | pointer to input vector 1 |
input_2_vect | const int16_t * | in | pointer to input vector 2 |
input_1_offset | const int32_t | in | offset for input 1 |
input_1_mult | const int32_t | in | multiplier for input 1 |
input_1_shift | const int32_t | in | shift for input 1 |
input_2_offset | const int32_t | in | offset for input 2 |
input_2_mult | const int32_t | in | multiplier for input 2 |
input_2_shift | const int32_t | in | shift for input 2 |
left_shift | const int32_t | in | Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
output | int16_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s8 elementwise multiplication of two tensors with support for broadcasting.
arm_cmsis_nn_status arm_mul_s8( const int8_t *input1_data, const cmsis_nn_dims *input1_dims, const int8_t *input2_data, const cmsis_nn_dims *input2_dims, const int32_t input1_offset, const int32_t input2_offset, int8_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max)s8 elementwise multiplication of two tensors with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input1_data | const int8_t * | in | pointer to input tensor 1 |
input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
input2_data | const int8_t * | in | pointer to input tensor 2 |
input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
input1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
output_data | int8_t * | out | pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
s8 elementwise multiplication of scalar and vector
arm_cmsis_nn_status arm_mul_scalar_s8( const int8_t *input_1_vect, const int8_t *input_2_vect, const int32_t input_1_offset, const int32_t input_2_offset, int8_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s8 elementwise multiplication of scalar and vector
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int8_t * | in | pointer to input scalar |
input_2_vect | const int8_t * | in | pointer to input vector |
input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
output | int8_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s8 elementwise multiplication
arm_cmsis_nn_status arm_elementwise_mul_s8( const int8_t *input_1_vect, const int8_t *input_2_vect, const int32_t input_1_offset, const int32_t input_2_offset, int8_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s8 elementwise multiplication
Supported framework: TensorFlow Lite micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int8_t * | in | pointer to input vector 1 |
input_2_vect | const int8_t * | in | pointer to input vector 2 |
input_1_offset | const int32_t | in | offset for input 1. Range: -127 to 128 |
input_2_offset | const int32_t | in | offset for input 2. Range: -127 to 128 |
output | int8_t * | in, out | pointer to output vector |
out_offset | const int32_t | in | output offset. Range: -128 to 127 |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -128 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 127 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
s16 elementwise multiplication of two tensors with support for broadcasting.
arm_cmsis_nn_status arm_mul_s16( const int16_t *input1_data, const cmsis_nn_dims *input1_dims, const int16_t *input2_data, const cmsis_nn_dims *input2_dims, const int32_t input1_offset, const int32_t input2_offset, int16_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max)s16 elementwise multiplication of two tensors with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input1_data | const int16_t * | in | pointer to input tensor 1 |
input1_dims | const cmsis_nn_dims * | in | pointer to input tensor 1 dimensions |
input2_data | const int16_t * | in | pointer to input tensor 2 |
input2_dims | const cmsis_nn_dims * | in | pointer to input tensor 2 dimensions |
input1_offset | const int32_t | in | offset for input 1. Not used. |
input2_offset | const int32_t | in | offset for input 2. Not used. |
output_data | int16_t * | out | pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | pointer to output tensor dimensions |
out_offset | const int32_t | in | output offset. Not used. |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. |
s16 elementwise multiplication of scalar and vector
arm_cmsis_nn_status arm_mul_scalar_s16( const int16_t *input_1_vect, const int16_t *input_2_vect, const int32_t input_1_offset, const int32_t input_2_offset, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s16 elementwise multiplication of scalar and vector
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int16_t * | in | pointer to input scalar |
input_2_vect | const int16_t * | in | pointer to input vector |
input_1_offset | const int32_t | in | offset for input 1. Not used. |
input_2_offset | const int32_t | in | offset for input 2. Not used. |
output | int16_t * | out | pointer to output vector |
out_offset | const int32_t | in | output offset. Not used. |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns ARM_CMSIS_NN_SUCCESS |
s16 elementwise multiplication
arm_cmsis_nn_status arm_elementwise_mul_s16( const int16_t *input_1_vect, const int16_t *input_2_vect, const int32_t input_1_offset, const int32_t input_2_offset, int16_t *output, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift, const int32_t out_activation_min, const int32_t out_activation_max, const int32_t block_size)s16 elementwise multiplication
Supported framework: TensorFlow Lite micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_1_vect | const int16_t * | in | pointer to input vector 1 |
input_2_vect | const int16_t * | in | pointer to input vector 2 |
input_1_offset | const int32_t | in | offset for input 1. Not used. |
input_2_offset | const int32_t | in | offset for input 2. Not used. |
output | int16_t * | in, out | pointer to output vector |
out_offset | const int32_t | in | output offset. Not used. |
out_mult | const int32_t | in | output multiplier |
out_shift | const int32_t | in | output shift |
out_activation_min | const int32_t | in | minimum value to clamp output to. Min: -32768 |
out_activation_max | const int32_t | in | maximum value to clamp output to. Max: 32767 |
block_size | const int32_t | in | number of samples |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
s8 elementwise minimum w/ support for broadcasting and scalar inputs.
arm_cmsis_nn_status arm_minimum_s8( const cmsis_nn_context *ctx, const int8_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int8_t *input_2_data, const cmsis_nn_dims *input_2_dims, int8_t *output_data, const cmsis_nn_dims *output_dims)s8 elementwise minimum w/ support for broadcasting and scalar inputs.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int8_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int8_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | int8_t * | out | Pointer to the output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
s8 elementwise maximum w/ support for broadcasting and scalar inputs.
arm_cmsis_nn_status arm_maximum_s8( const cmsis_nn_context *ctx, const int8_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int8_t *input_2_data, const cmsis_nn_dims *input_2_dims, int8_t *output_data, const cmsis_nn_dims *output_dims)s8 elementwise maximum w/ support for broadcasting and scalar inputs.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int8_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int8_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | int8_t * | out | Pointer to the output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
s16 elementwise minimum w/ support for broadcasting and scalar inputs.
arm_cmsis_nn_status arm_minimum_s16( const cmsis_nn_context *ctx, const int16_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int16_t *input_2_data, const cmsis_nn_dims *input_2_dims, int16_t *output_data, const cmsis_nn_dims *output_dims)s16 elementwise minimum w/ support for broadcasting and scalar inputs.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int16_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int16_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | int16_t * | out | Pointer to the output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
s16 elementwise maximum w/ support for broadcasting and scalar inputs.
arm_cmsis_nn_status arm_maximum_s16( const cmsis_nn_context *ctx, const int16_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int16_t *input_2_data, const cmsis_nn_dims *input_2_dims, int16_t *output_data, const cmsis_nn_dims *output_dims)s16 elementwise maximum w/ support for broadcasting and scalar inputs.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int16_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int16_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | int16_t * | out | Pointer to the output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
s8 elementwise comparison with support for broadcasting.
arm_cmsis_nn_status arm_comparison_s8( const cmsis_nn_context *ctx, const int8_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int8_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, arm_nn_compare_operation operation)s8 elementwise comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int8_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int8_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
operation | arm_nn_compare_operation | in | Comparison operation to perform |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
s16 elementwise comparison with support for broadcasting.
arm_cmsis_nn_status arm_comparison_s16( const cmsis_nn_context *ctx, const int16_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int16_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift, arm_nn_compare_operation operation)s16 elementwise comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int16_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int16_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
operation | arm_nn_compare_operation | in | Comparison operation to perform |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL. |
s8 elementwise equality comparison with support for broadcasting.
arm_cmsis_nn_status arm_equal_s8( const cmsis_nn_context *ctx, const int8_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int8_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s8 elementwise equality comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int8_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int8_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s8 elementwise inequality comparison with support for broadcasting.
arm_cmsis_nn_status arm_not_equal_s8( const cmsis_nn_context *ctx, const int8_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int8_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s8 elementwise inequality comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int8_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int8_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s8 elementwise greater-than comparison with support for broadcasting.
arm_cmsis_nn_status arm_greater_s8( const cmsis_nn_context *ctx, const int8_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int8_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s8 elementwise greater-than comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int8_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int8_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s8 elementwise greater-or-equal comparison with support for broadcasting.
arm_cmsis_nn_status arm_greater_equal_s8( const cmsis_nn_context *ctx, const int8_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int8_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s8 elementwise greater-or-equal comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int8_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int8_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s8 elementwise less-than comparison with support for broadcasting.
arm_cmsis_nn_status arm_less_s8( const cmsis_nn_context *ctx, const int8_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int8_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s8 elementwise less-than comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int8_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int8_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s8 elementwise less-or-equal comparison with support for broadcasting.
arm_cmsis_nn_status arm_less_equal_s8( const cmsis_nn_context *ctx, const int8_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int8_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s8 elementwise less-or-equal comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int8_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int8_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s16 elementwise equality comparison with support for broadcasting.
arm_cmsis_nn_status arm_equal_s16( const cmsis_nn_context *ctx, const int16_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int16_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s16 elementwise equality comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int16_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int16_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s16 elementwise inequality comparison with support for broadcasting.
arm_cmsis_nn_status arm_not_equal_s16( const cmsis_nn_context *ctx, const int16_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int16_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s16 elementwise inequality comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int16_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int16_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s16 elementwise greater-than comparison with support for broadcasting.
arm_cmsis_nn_status arm_greater_s16( const cmsis_nn_context *ctx, const int16_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int16_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s16 elementwise greater-than comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int16_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int16_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s16 elementwise greater-or-equal comparison with support for broadcasting.
arm_cmsis_nn_status arm_greater_equal_s16( const cmsis_nn_context *ctx, const int16_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int16_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s16 elementwise greater-or-equal comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int16_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int16_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s16 elementwise less-than comparison with support for broadcasting.
arm_cmsis_nn_status arm_less_s16( const cmsis_nn_context *ctx, const int16_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int16_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s16 elementwise less-than comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int16_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int16_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
s16 elementwise less-or-equal comparison with support for broadcasting.
arm_cmsis_nn_status arm_less_equal_s16( const cmsis_nn_context *ctx, const int16_t *input_1_data, const cmsis_nn_dims *input_1_dims, const int16_t *input_2_data, const cmsis_nn_dims *input_2_dims, bool *output_data, const cmsis_nn_dims *output_dims, const int32_t input_1_offset, const int32_t input_1_mult, const int32_t input_1_shift, const int32_t input_2_offset, const int32_t input_2_mult, const int32_t input_2_shift, const int32_t left_shift)s16 elementwise less-or-equal comparison with support for broadcasting.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused; may be NULL. |
input_1_data | const int16_t * | in | Pointer to input1 tensor |
input_1_dims | const cmsis_nn_dims * | in | Input1 tensor dimensions |
input_2_data | const int16_t * | in | Pointer to input2 tensor |
input_2_dims | const cmsis_nn_dims * | in | Input2 tensor dimensions |
output_data | bool * | out | Pointer to the output tensor (bool values) |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
input_1_offset | const int32_t | in | Zero-point for input1 tensor |
input_1_mult | const int32_t | in | Multiplier for input1 tensor |
input_1_shift | const int32_t | in | Shift for input1 tensor |
input_2_offset | const int32_t | in | Zero-point for input2 tensor |
input_2_mult | const int32_t | in | Multiplier for input2 tensor |
input_2_shift | const int32_t | in | Shift for input2 tensor |
left_shift | const int32_t | in | Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel. |
Returns
| Description |
|---|
| As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments. |
Q7 RELU function.
void arm_relu_q7(int8_t *data, uint16_t size)Q7 RELU function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
data | int8_t * | in, out | pointer to input |
size | uint16_t | in | number of elements |
Q7 RELU6 function.
void arm_relu6_q7(int8_t *data, uint16_t size)Q7 RELU6 function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
data | int8_t * | in, out | pointer to input |
size | uint16_t | in | number of elements |
Q15 RELU function.
void arm_relu_q15(int16_t *data, uint16_t size)Q15 RELU function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
data | int16_t * | in, out | pointer to input |
size | uint16_t | in | number of elements |
S8 clamp function.
arm_cmsis_nn_status arm_clamp_s8( const int8_t *input, const int8_t act_min, const int8_t act_max, int8_t *output, const int32_t output_size)S8 clamp function.
This function clamps each element in the input tensor to the range. This can be useful for activations such as relu(0, 6), relu(-1,1), etc.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input |
act_min | const int8_t | in | Minimum value to clamp to |
act_max | const int8_t | in | Maximum value to clamp to |
output | int8_t * | out | Pointer to output |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
S16 clamp function.
arm_cmsis_nn_status arm_clamp_s16( const int16_t *input, const int16_t act_min, const int16_t act_max, int16_t *output, const int32_t output_size)S16 clamp function.
This function clamps each element in the input tensor to the range. This can be useful for activations such as relu(0, 6), relu(-1,1), etc.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to input |
act_min | const int16_t | in | Minimum value to clamp to |
act_max | const int16_t | in | Maximum value to clamp to |
output | int16_t * | out | Pointer to output |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127.
arm_cmsis_nn_status arm_relu_s8( const int8_t *input, const int32_t input_offset, const int32_t output_offset, const int32_t output_multiplier, const int32_t output_shift, int8_t *output, const int32_t output_size)S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input buffer |
input_offset | const int32_t | in | Input tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier | const int32_t | in | Output multiplier |
output_shift | const int32_t | in | Output shift |
output | int8_t * | out | Pointer to the output buffer |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.
arm_cmsis_nn_status arm_relu_generic_s8( const int8_t *input, const int32_t input_offset, const int32_t output_offset, const int32_t output_multiplier, const int32_t output_shift, const int32_t act_min, const int32_t act_max, int8_t *output, const int32_t output_size)S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input buffer |
input_offset | const int32_t | in | Input tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier | const int32_t | in | Output multiplier |
output_shift | const int32_t | in | Output shift |
act_min | const int32_t | in | Minimum value to clamp the output to |
act_max | const int32_t | in | Maximum value to clamp the output to |
output | int8_t * | out | Pointer to the output buffer |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767.
arm_cmsis_nn_status arm_relu_s16( const int16_t *input, const int32_t input_offset, const int32_t output_offset, const int32_t output_multiplier, const int32_t output_shift, int16_t *output, const int32_t output_size)S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input buffer |
input_offset | const int32_t | in | Input tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier | const int32_t | in | Output multiplier |
output_shift | const int32_t | in | Output shift |
output | int16_t * | out | Pointer to the output buffer |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.
arm_cmsis_nn_status arm_relu_generic_s16( const int16_t *input, const int32_t input_offset, const int32_t output_offset, const int32_t output_multiplier, const int32_t output_shift, const int32_t act_min, const int32_t act_max, int16_t *output, const int32_t output_size)S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input buffer |
input_offset | const int32_t | in | Input tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier | const int32_t | in | Output multiplier |
output_shift | const int32_t | in | Output shift |
act_min | const int32_t | in | Minimum value to clamp the output to |
act_max | const int32_t | in | Maximum value to clamp the output to |
output | int16_t * | out | Pointer to the output buffer |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
S8 Leaky ReLU activation function.
arm_cmsis_nn_status arm_leaky_relu_s8( const int8_t *input, const int32_t input_offset, const int32_t output_offset, const int32_t output_multiplier_alpha, const int32_t output_shift_alpha, const int32_t output_multiplier_identity, const int32_t output_shift_identity, int8_t *output, const int32_t output_size)S8 Leaky ReLU activation function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input buffer |
input_offset | const int32_t | in | Input tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier_alpha | const int32_t | in | Output multiplier for the alpha parameter |
output_shift_alpha | const int32_t | in | Output shift for the alpha parameter |
output_multiplier_identity | const int32_t | in | Output multiplier for the identity parameter |
output_shift_identity | const int32_t | in | Output shift for the identity parameter |
output | int8_t * | out | Pointer to the output buffer |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
S16 Leaky ReLU activation function.
arm_cmsis_nn_status arm_leaky_relu_s16( const int16_t *input, const int32_t input_offset, const int32_t output_offset, const int32_t output_multiplier_alpha, const int32_t output_shift_alpha, const int32_t output_multiplier_identity, const int32_t output_shift_identity, int16_t *output, const int32_t output_size)S16 Leaky ReLU activation function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input buffer |
input_offset | const int32_t | in | Input tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier_alpha | const int32_t | in | Output multiplier for the alpha parameter |
output_shift_alpha | const int32_t | in | Output shift for the alpha parameter |
output_multiplier_identity | const int32_t | in | Output multiplier for the identity parameter |
output_shift_identity | const int32_t | in | Output shift for the identity parameter |
output | int16_t * | out | Pointer to the output buffer |
output_size | const int32_t | in | Number of elements in the input tensor |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
Logistic activation function for s16.
arm_cmsis_nn_status arm_logistic_s16( const int16_t *input, int16_t *output, const int32_t input_size, int32_t input_multiplier, int32_t input_left_shift)Logistic activation function for s16.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input tensor |
output | int16_t * | out | Pointer to the output tensor |
input_size | const int32_t | in | Number of elements in the input tensor |
input_multiplier | int32_t | in | Input quantization multiplier |
input_left_shift | int32_t | in | Input quantization shift within the range [0, 31] |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation |
Tanh activation function for s16.
arm_cmsis_nn_status arm_tanh_s16( const int16_t *input, int16_t *output, const int32_t input_size, int32_t input_multiplier, int32_t input_left_shift)Tanh activation function for s16.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input tensor |
output | int16_t * | out | Pointer to the output tensor |
input_size | const int32_t | in | Number of elements in the input tensor |
input_multiplier | int32_t | in | Input quantization multiplier |
input_left_shift | int32_t | in | Input quantization shift within the range [0, 31] |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation |
s16 neural network activation function using direct table look-up
arm_cmsis_nn_status arm_nn_activation_s16( const int16_t *input, int16_t *output, const int32_t size, const int32_t left_shift, const arm_nn_activation_type type)s16 neural network activation function using direct table look-up
Supported framework: TensorFlow Lite for Microcontrollers. This activation function must be bit precise congruent with the corresponding TFLM tanh and sigmoid activation functions
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | pointer to input data |
output | int16_t * | out | pointer to output |
size | const int32_t | in | number of elements |
left_shift | const int32_t | in | bit-width of the integer part, assumed to be smaller than 3. |
type | const arm_nn_activation_type | in | type of activation functions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
S8 Hard-Swish activation function (compatibility version).
arm_cmsis_nn_status arm_hard_swish_compat_s8( const int8_t *input, const int32_t input_offset, const int32_t output_offset, const int32_t output_multiplier_fp, const int32_t output_multiplier_exp, const int32_t relu_multiplier_fp, const int32_t relu_multiplier_exp, int8_t *output, const int32_t output_size)S8 Hard-Swish activation function (compatibility version).
This version is compatible with TFLite implementation of Hard-Swish. hires_input_scale = (1.0 / 128.0) * float(input_scale) relu_scale = 3.0 / 32768.0 out_mul_real = hires_input_scale / float(output_scale) relu_mul_real = hires_input_scale / relu_scale output_multiplier_fp, output_multiplier_exp = to_q15_exp(out_mul_real) relu_multiplier_fp, relu_multiplier_exp = to_q15_exp(relu_mul_real) Here to_q15_exp quantizes to Q31 with a frexp exponent, then rounds and saturates the Q31 multiplier to Q15. For input_scale = output_scale = 0.125, the output pair is (16384, -6) and the ReLU pair is (21845, 4).
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input buffer |
input_offset | const int32_t | in | Input tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier_fp | const int32_t | in | Output multiplier in fixed point format |
output_multiplier_exp | const int32_t | in | Exponent for output multiplier |
relu_multiplier_fp | const int32_t | in | ReLU6 multiplier in fixed point format |
relu_multiplier_exp | const int32_t | in | Exponent for ReLU6 multiplier |
output | int8_t * | out | Pointer to the output buffer |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR if output_multiplier_exp is positive. |
S8 Hard-Swish activation function (precise version).
arm_cmsis_nn_status arm_hard_swish_precise_s8( const int8_t *input, const int32_t input_offset, const int32_t output_offset, const int32_t output_multiplier, const int32_t output_shift, const int32_t relu_q3, const int32_t relu_q6, const int32_t prescale, int8_t *output, const int32_t output_size)S8 Hard-Swish activation function (precise version).
This version uses int32_t for intermediate computations to provide better accuracy. relu_q3, relu_q6 = round(3 / input_scale), round(6 / input_scale) prescale = min value to multiply input by to avoid overflow in intermediate computations M = ((input_scale**2) / (6.0 * output_scale)) * (1 << prescale) output_multiplier, output_shift = quantize_multiplier(M)
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input buffer |
input_offset | const int32_t | in | Input tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier | const int32_t | in | Output multiplier |
output_shift | const int32_t | in | Output shift |
relu_q3 | const int32_t | in | ReLU6 Q3 value |
relu_q6 | const int32_t | in | ReLU6 Q6 value |
prescale | const int32_t | in | Prescale to apply to input |
output | int8_t * | out | Pointer to the output buffer |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
S16 Hard-Swish activation function (precise version).
arm_cmsis_nn_status arm_hard_swish_precise_s16( const int16_t *input, const int32_t input_offset, const int32_t output_offset, const int32_t output_multiplier, const int32_t output_shift, const int32_t relu_q3, const int32_t relu_q6, const int32_t prescale, int16_t *output, const int32_t output_size)S16 Hard-Swish activation function (precise version).
This version uses int32_t for intermediate computations to provide better accuracy. relu_q3, relu_q6 = round(3 / input_scale), round(6 / input_scale) prescale = min value to multiply input by to avoid overflow in intermediate computations M = ((input_scale**2) / (6.0 * output_scale)) * (1 << prescale) output_multiplier, output_shift = quantize_multiplier(M)
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input buffer |
input_offset | const int32_t | in | Input tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier | const int32_t | in | Output multiplier |
output_shift | const int32_t | in | Output shift |
relu_q3 | const int32_t | in | ReLU6 Q3 value |
relu_q6 | const int32_t | in | ReLU6 Q6 value |
prescale | const int32_t | in | Prescale to apply to input |
output | int16_t * | out | Pointer to the output buffer |
output_size | const int32_t | in | Number of elements in the tensor |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
S8 PReLU activation function.
arm_cmsis_nn_status arm_prelu_s8( const cmsis_nn_dims *input_dims, const int8_t *input, const cmsis_nn_dims *alpha_dims, const int8_t *alpha, const int32_t input_offset, const int32_t alpha_offset, const int32_t output_offset, const int32_t output_multiplier_identity, const int32_t output_shift_identity, const int32_t output_multiplier_alpha, const int32_t output_shift_alpha, const cmsis_nn_dims *output_dims, int8_t *output)S8 PReLU activation function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input | const int8_t * | in | Pointer to the input buffer |
alpha_dims | const cmsis_nn_dims * | in | Alpha tensor dimensions. Format: [N, H, W, C] |
alpha | const int8_t * | in | Pointer to the alpha buffer |
input_offset | const int32_t | in | Input tensor zero offset |
alpha_offset | const int32_t | in | Alpha tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier_identity | const int32_t | in | Output multiplier 1 |
output_shift_identity | const int32_t | in | Output shift 1 |
output_multiplier_alpha | const int32_t | in | Output multiplier 2 |
output_shift_alpha | const int32_t | in | Output shift 2 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output | int8_t * | out | Pointer to the output buffer |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, alpha does not broadcast into the input, or the output dimensions do not match the input dimensions. |
Elementwise S8 PReLU activation function.
arm_cmsis_nn_status arm_elementwise_prelu_s8( const int8_t *input, const int8_t *alpha, const int32_t input_offset, const int32_t alpha_offset, const int32_t out_offset, const int32_t output_multiplier_identity, const int32_t output_shift_identity, const int32_t output_multiplier_alpha, const int32_t output_shift_alpha, int8_t *output, const int32_t block_size)Elementwise S8 PReLU activation function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input buffer |
alpha | const int8_t * | in | Pointer to the alpha buffer (same shape as input) |
input_offset | const int32_t | in | Input tensor zero offset |
alpha_offset | const int32_t | in | Alpha tensor zero offset |
out_offset | const int32_t | in | Output tensor zero offset |
output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
output_shift_identity | const int32_t | in | Output shift when input >= 0 |
output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
output_shift_alpha | const int32_t | in | Output shift when input < 0 |
output | int8_t * | out | Pointer to the output buffer |
block_size | const int32_t | in | Number of elements to process |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
Scalar S8 PReLU activation function.
arm_cmsis_nn_status arm_prelu_scalar_s8( const int8_t *scalar_vect, const int8_t *non_scalar_vect, const bool scalar_is_input, const int32_t input_offset, const int32_t alpha_offset, const int32_t output_offset, const int32_t output_multiplier_identity, const int32_t output_shift_identity, const int32_t output_multiplier_alpha, const int32_t output_shift_alpha, int8_t *output, const int32_t block_size)Scalar S8 PReLU activation function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
scalar_vect | const int8_t * | in | Pointer to the scalar buffer (single value) |
non_scalar_vect | const int8_t * | in | Pointer to the non-scalar buffer |
scalar_is_input | const bool | in | True if the scalar buffer holds the input value, false if it holds alpha |
input_offset | const int32_t | in | Input tensor zero offset |
alpha_offset | const int32_t | in | Alpha tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
output_shift_identity | const int32_t | in | Output shift when input >= 0 |
output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
output_shift_alpha | const int32_t | in | Output shift when input < 0 |
output | int8_t * | out | Pointer to the output buffer |
block_size | const int32_t | in | Number of elements to process when the non-scalar vector is used |
Returns
| Description |
|---|
| The function returns ARM_MATH_SUCCESS |
S16 PReLU activation function.
arm_cmsis_nn_status arm_prelu_s16( const cmsis_nn_dims *input_dims, const int16_t *input, const cmsis_nn_dims *alpha_dims, const int16_t *alpha, const int32_t input_offset, const int32_t alpha_offset, const int32_t output_offset, const int32_t output_multiplier_identity, const int32_t output_shift_identity, const int32_t output_multiplier_alpha, const int32_t output_shift_alpha, const cmsis_nn_dims *output_dims, int16_t *output)S16 PReLU activation function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
input | const int16_t * | in | Pointer to the input buffer |
alpha_dims | const cmsis_nn_dims * | in | Alpha tensor dimensions. Format: [N, H, W, C] |
alpha | const int16_t * | in | Pointer to the alpha buffer |
input_offset | const int32_t | in | Input tensor zero offset |
alpha_offset | const int32_t | in | Alpha tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
output_shift_identity | const int32_t | in | Output shift when input >= 0 |
output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
output_shift_alpha | const int32_t | in | Output shift when input < 0 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C_OUT] |
output | int16_t * | out | Pointer to the output buffer |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, alpha does not broadcast into the input, or the output dimensions do not match the input dimensions. |
Elementwise S16 PReLU activation function.
arm_cmsis_nn_status arm_elementwise_prelu_s16( const int16_t *input, const int16_t *alpha, const int32_t input_offset, const int32_t alpha_offset, const int32_t out_offset, const int32_t output_multiplier_identity, const int32_t output_shift_identity, const int32_t output_multiplier_alpha, const int32_t output_shift_alpha, int16_t *output, const int32_t block_size)Elementwise S16 PReLU activation function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input buffer |
alpha | const int16_t * | in | Pointer to the alpha buffer (same shape as input) |
input_offset | const int32_t | in | Input tensor zero offset |
alpha_offset | const int32_t | in | Alpha tensor zero offset |
out_offset | const int32_t | in | Output tensor zero offset |
output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
output_shift_identity | const int32_t | in | Output shift when input >= 0 |
output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
output_shift_alpha | const int32_t | in | Output shift when input < 0 |
output | int16_t * | out | Pointer to the output buffer |
block_size | const int32_t | in | Number of elements to process |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS |
Scalar S16 PReLU activation function.
arm_cmsis_nn_status arm_prelu_scalar_s16( const int16_t *scalar_vect, const int16_t *non_scalar_vect, const bool scalar_is_input, const int32_t input_offset, const int32_t alpha_offset, const int32_t output_offset, const int32_t output_multiplier_identity, const int32_t output_shift_identity, const int32_t output_multiplier_alpha, const int32_t output_shift_alpha, int16_t *output, const int32_t block_size)Scalar S16 PReLU activation function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
scalar_vect | const int16_t * | in | Pointer to the scalar buffer (single value) |
non_scalar_vect | const int16_t * | in | Pointer to the non-scalar buffer |
scalar_is_input | const bool | in | True if the scalar buffer holds the input value, false if it holds alpha |
input_offset | const int32_t | in | Input tensor zero offset |
alpha_offset | const int32_t | in | Alpha tensor zero offset |
output_offset | const int32_t | in | Output tensor zero offset |
output_multiplier_identity | const int32_t | in | Output multiplier when input >= 0 |
output_shift_identity | const int32_t | in | Output shift when input >= 0 |
output_multiplier_alpha | const int32_t | in | Output multiplier when input < 0 |
output_shift_alpha | const int32_t | in | Output shift when input < 0 |
output | int16_t * | out | Pointer to the output buffer |
block_size | const int32_t | in | Number of elements to process when the non-scalar vector is used |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS |
s8 average pooling function.
arm_cmsis_nn_status arm_avgpool_s8( const cmsis_nn_context *ctx, const cmsis_nn_pool_params *pool_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 average pooling function.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context. Size ctx->buf with arm_avgpool_s8_get_buffer_size(output_dims->w, input_dims->c). Note that it takes the output width and the input channel count rather than a `cmsis_nn_dims`. `arm_avgpool_s8_get_buffer_size_dsp()` and `arm_avgpool_s8_get_buffer_size_mve()` size the same buffer for a specific target. It returns 0 where no buffer is needed, in which case ctx->buf may be NULL. The caller is expected to clear the buffer, if applicable, for security reasons. |
pool_params | const cmsis_nn_pool_params * | in | Pooling parameters |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [H, W] Argument N and C are not used. |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN. |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns (an extent of 0 or less), which writes nothing and does not use ctx `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments, including a negative channel count, a pooling window that does not overlap the input, window positions (output index times stride minus padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t, or, on builds without MVE, a NULL ctx, or a NULL ctx->buf where the sizer asks for a buffer. Nothing is written to output_data then. |
Get the required buffer size for S8 average pooling function.
int32_t arm_avgpool_s8_get_buffer_size(const int dim_dst_width, const int ch_src)Get the required buffer size for S8 average pooling function.
Unlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. arm_avgpool_s8() depends on that sentinel being non-zero, since it reads a non-zero size as “ctx->buf is required” before touching the accumulator buffer.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dim_dst_width | const int | in | output tensor dimension |
ch_src | const int | in | number of input tensor channels |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Get the required buffer size for S8 average pooling function for processors with DSP extension.
int32_t arm_avgpool_s8_get_buffer_size_dsp(const int dim_dst_width, const int ch_src)Get the required buffer size for S8 average pooling function for processors with DSP extension.
Unlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. arm_avgpool_s8() depends on that sentinel being non-zero, since it reads a non-zero size as “ctx->buf is required” before touching the accumulator buffer.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dim_dst_width | const int | in | output tensor dimension |
ch_src | const int | in | number of input tensor channels |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case.
int32_t arm_avgpool_s8_get_buffer_size_mve(const int dim_dst_width, const int ch_src)Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case.
Unlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. arm_avgpool_s8() depends on that sentinel being non-zero, since it reads a non-zero size as “ctx->buf is required” before touching the accumulator buffer.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dim_dst_width | const int | in | output tensor dimension |
ch_src | const int | in | number of input tensor channels |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
s16 average pooling function.
arm_cmsis_nn_status arm_avgpool_s16( const cmsis_nn_context *ctx, const cmsis_nn_pool_params *pool_params, const cmsis_nn_dims *input_dims, const int16_t *input_data, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims, int16_t *output_data)s16 average pooling function.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Function context. Size ctx->buf with arm_avgpool_s16_get_buffer_size(output_dims->w, input_dims->c). Note that it takes the output width and the input channel count rather than a `cmsis_nn_dims`. `arm_avgpool_s16_get_buffer_size_dsp()` and `arm_avgpool_s16_get_buffer_size_mve()` size the same buffer for a specific target. It returns 0 where no buffer is needed, in which case ctx->buf may be NULL. The caller is expected to clear the buffer, if applicable, for security reasons. |
pool_params | const cmsis_nn_pool_params * | in | Pooling parameters |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] |
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [H, W] Argument N and C are not used. |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN. |
output_data | int16_t * | out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns (an extent of 0 or less), which writes nothing and does not use ctx `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments, including a negative channel count, a pooling window that does not overlap the input, window positions (output index times stride minus padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t, or, on builds that use the buffer, a NULL ctx, or a NULL ctx->buf where the sizer asks for a buffer. Nothing is written to output_data then. |
Get the required buffer size for S16 average pooling function.
int32_t arm_avgpool_s16_get_buffer_size(const int dim_dst_width, const int ch_src)Get the required buffer size for S16 average pooling function.
As in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dim_dst_width | const int | in | output tensor dimension |
ch_src | const int | in | number of input tensor channels |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Get the required buffer size for S16 average pooling function for processors with DSP extension.
int32_t arm_avgpool_s16_get_buffer_size_dsp(const int dim_dst_width, const int ch_src)Get the required buffer size for S16 average pooling function for processors with DSP extension.
As in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dim_dst_width | const int | in | output tensor dimension |
ch_src | const int | in | number of input tensor channels |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case.
int32_t arm_avgpool_s16_get_buffer_size_mve(const int dim_dst_width, const int ch_src)Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case.
As in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
dim_dst_width | const int | in | output tensor dimension |
ch_src | const int | in | number of input tensor channels |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t |
s8 max pooling function.
arm_cmsis_nn_status arm_max_pool_s8( const cmsis_nn_context *ctx, const cmsis_nn_pool_params *pool_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 max pooling function.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_max_pool_s8_get_buffer_size(). Max pooling needs no accumulator scratch, unlike `arm_avgpool_s8()`, whose sizer does not describe this argument. |
pool_params | const cmsis_nn_pool_params * | in | Pooling parameters |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] |
input_data | const int8_t * | in | Input (activation) data pointer. The input tensor must not overlap with the output tensor. Data type: int8 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [H, W] Argument N and C are not used. |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN. |
output_data | int8_t * | out | Output data pointer. Data type: int8 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns, which writes nothing `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments: a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to output_data then. |
s16 max pooling function.
arm_cmsis_nn_status arm_max_pool_s16( const cmsis_nn_context *ctx, const cmsis_nn_pool_params *pool_params, const cmsis_nn_dims *input_dims, const int16_t *src, const cmsis_nn_dims *filter_dims, const cmsis_nn_dims *output_dims, int16_t *dst)s16 max pooling function.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_max_pool_s16_get_buffer_size(). Max pooling needs no accumulator scratch, unlike `arm_avgpool_s16()`, whose sizer does not describe this argument. |
pool_params | const cmsis_nn_pool_params * | in | Pooling parameters |
input_dims | const cmsis_nn_dims * | in | Input (activation) tensor dimensions. Format: [H, W, C_IN] |
src | const int16_t * | in | Input (activation) data pointer. The input tensor must not overlap with the output tensor. Data type: int16 |
filter_dims | const cmsis_nn_dims * | in | Filter tensor dimensions. Format: [H, W] Argument N and C are not used. |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN. |
dst | int16_t * | in, out | Output data pointer. Data type: int16 |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns, which writes nothing `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments: a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then. |
S8 softmax function.
void arm_softmax_s8( const int8_t *input, const int32_t num_rows, const int32_t row_size, const int32_t mult, const int32_t shift, const int32_t diff_min, int8_t *output)S8 softmax function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input tensor |
num_rows | const int32_t | in | Number of rows in the input tensor |
row_size | const int32_t | in | Number of elements in each input row |
mult | const int32_t | in | Input quantization multiplier |
shift | const int32_t | in | Input quantization shift within the range [0, 31] |
diff_min | const int32_t | in | Minimum difference with max in row. Used to check if the quantized exponential operation can be performed |
output | int8_t * | out | Pointer to the output tensor |
S8 to s16 softmax function.
void arm_softmax_s8_s16( const int8_t *input, const int32_t num_rows, const int32_t row_size, const int32_t mult, const int32_t shift, const int32_t diff_min, int16_t *output)S8 to s16 softmax function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input tensor |
num_rows | const int32_t | in | Number of rows in the input tensor |
row_size | const int32_t | in | Number of elements in each input row |
mult | const int32_t | in | Input quantization multiplier |
shift | const int32_t | in | Input quantization shift within the range [0, 31] |
diff_min | const int32_t | in | Minimum difference with max in row. Used to check if the quantized exponential operation can be performed |
output | int16_t * | out | Pointer to the output tensor |
S16 softmax function.
arm_cmsis_nn_status arm_softmax_s16( const int16_t *input, const int32_t num_rows, const int32_t row_size, const int32_t mult, const int32_t shift, const cmsis_nn_softmax_lut_s16 *softmax_params, int16_t *output)S16 softmax function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input tensor |
num_rows | const int32_t | in | Number of rows in the input tensor |
row_size | const int32_t | in | Number of elements in each input row |
mult | const int32_t | in | Input quantization multiplier |
shift | const int32_t | in | Input quantization shift within the range [0, 31] |
softmax_params | const cmsis_nn_softmax_lut_s16 * | in | Softmax s16 layer parameters with two pointers to LUTs speficied below. For indexing the high 9 bits are used and 7 remaining for interpolation. That means 512 entries for the 9-bit indexing and 1 extra for interpolation, i.e. 513 values for each LUT. - Lookup table for exp(x), where x uniform distributed between [-10.0 , 0.0] - Lookup table for 1 / (1 + x), where x uniform distributed between [0.0 , 1.0] |
output | int16_t * | out | Pointer to the output tensor |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation |
U8 softmax function.
void arm_softmax_u8( const uint8_t *input, const int32_t num_rows, const int32_t row_size, const int32_t mult, const int32_t shift, const int32_t diff_min, uint8_t *output)U8 softmax function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const uint8_t * | in | Pointer to the input tensor |
num_rows | const int32_t | in | Number of rows in the input tensor |
row_size | const int32_t | in | Number of elements in each input row |
mult | const int32_t | in | Input quantization multiplier |
shift | const int32_t | in | Input quantization shift within the range [0, 31] |
diff_min | const int32_t | in | Minimum difference with max in row. Used to check if the quantized exponential operation can be performed |
output | uint8_t * | out | Pointer to the output tensor |
Reshape a s8 vector into another with different shape.
void arm_reshape_s8(const int8_t *input, int8_t *output, const uint32_t total_size)Reshape a s8 vector into another with different shape.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | points to the s8 input vector |
output | int8_t * | out | points to the s8 output vector |
total_size | const uint32_t | in | total size of the input and output vectors in bytes |
Nearest neighbor resize function for s8 data.
arm_cmsis_nn_status arm_resize_nearest_neighbor_s8( const cmsis_nn_context *ctx, const cmsis_nn_resize_params *resize_params, const cmsis_nn_dims *input_shape, const int8_t *input_data, const cmsis_nn_dims *output_size_shape, const int32_t *output_size_data, const cmsis_nn_dims *output_shape, int8_t *output_data)Nearest neighbor resize function for s8 data.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Pointer to the context buffer. The buffer must hold at least (output_height + output_width) int32_t elements. |
resize_params | const cmsis_nn_resize_params * | in | Resize parameters |
input_shape | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C] |
input_data | const int8_t * | in | Pointer to input tensor data |
output_size_shape | const cmsis_nn_dims * | in | Output size tensor dimensions |
output_size_data | const int32_t * | in | Output size tensor data |
output_shape | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C] |
output_data | int8_t * | out | Pointer to output tensor data |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Nearest neighbor resize function for s16 data.
arm_cmsis_nn_status arm_resize_nearest_neighbor_s16( const cmsis_nn_context *ctx, const cmsis_nn_resize_params *resize_params, const cmsis_nn_dims *input_shape, const int16_t *input_data, const cmsis_nn_dims *output_size_shape, const int32_t *output_size_data, const cmsis_nn_dims *output_shape, int16_t *output_data)Nearest neighbor resize function for s16 data.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Pointer to the context buffer. The buffer must hold at least (output_height + output_width) int32_t elements. |
resize_params | const cmsis_nn_resize_params * | in | Resize parameters |
input_shape | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C] |
input_data | const int16_t * | in | Pointer to input tensor data |
output_size_shape | const cmsis_nn_dims * | in | Output size tensor dimensions |
output_size_data | const int32_t * | in | Output size tensor data |
output_shape | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H, W, C] |
output_data | int16_t * | out | Pointer to output tensor data |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Space to Depth function for s8 data type.
arm_cmsis_nn_status arm_space_to_depth_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const int32_t block_size, int8_t *output_data, const cmsis_nn_dims *output_dims)Space to Depth function for s8 data type.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to the input tensor. Data type: int8 |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
block_size | const int32_t | in | Block size for space to depth transformation |
output_data | int8_t * | out | Pointer to the output tensor. Data type: int8 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H/block_size, W/block_size, C_IN*block_size*block_size] |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Space to Depth function for s16 data type.
arm_cmsis_nn_status arm_space_to_depth_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const int32_t block_size, int16_t *output_data, const cmsis_nn_dims *output_dims)Space to Depth function for s16 data type.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to the input tensor. Data type: int16 |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
block_size | const int32_t | in | Block size for space to depth transformation |
output_data | int16_t * | out | Pointer to the output tensor. Data type: int16 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H/block_size, W/block_size, C_IN*block_size*block_size] |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Depth to Space function for s8 data type.
arm_cmsis_nn_status arm_depth_to_space_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const int32_t block_size, int8_t *output_data, const cmsis_nn_dims *output_dims)Depth to Space function for s8 data type.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to the input tensor. Data type: int8 |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
block_size | const int32_t | in | Block size for depth to space transformation |
output_data | int8_t * | out | Pointer to the output tensor. Data type: int8 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H*block_size, W*block_size, C_IN/(block_size*block_size)] |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Depth to Space function for s16 data type.
arm_cmsis_nn_status arm_depth_to_space_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const int32_t block_size, int16_t *output_data, const cmsis_nn_dims *output_dims)Depth to Space function for s16 data type.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to the input tensor. Data type: int16 |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
block_size | const int32_t | in | Block size for depth to space transformation |
output_data | int16_t * | out | Pointer to the output tensor. Data type: int16 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N, H*block_size, W*block_size, C_IN/(block_size*block_size)] |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Space to Batch ND function for s8 data type.
arm_cmsis_nn_status arm_space_to_batch_nd_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const cmsis_nn_tile *block_shape, const cmsis_nn_dims *pad, int8_t *output_data, const cmsis_nn_dims *output_dims, const int32_t output_offset)Space to Batch ND function for s8 data type.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to the input tensor. Data type: int8 |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
block_shape | const cmsis_nn_tile * | in | Block shape for space to batch transformation |
pad | const cmsis_nn_dims * | in | Padding for height and width. Format: [n->top, h->left, w->bottom, c->right] |
output_data | int8_t * | out | Pointer to the output tensor. Data type: int8 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N*block_shape[0]*block_shape[1], (H + pad_top + pad_bottom)/block_shape[0], (W + pad_left + pad_right)/block_shape[1], C_IN] |
output_offset | const int32_t | in | Zero offset for the output tensor |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Space to Batch ND function for s16 data type.
arm_cmsis_nn_status arm_space_to_batch_nd_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const cmsis_nn_tile *block_shape, const cmsis_nn_dims *pad, int16_t *output_data, const cmsis_nn_dims *output_dims, const int32_t output_offset)Space to Batch ND function for s16 data type.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to the input tensor. Data type: int16 |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
block_shape | const cmsis_nn_tile * | in | Block shape for space to batch transformation |
pad | const cmsis_nn_dims * | in | Padding for height and width. Format: [n->top, h->left, w->bottom, c->right] |
output_data | int16_t * | out | Pointer to the output tensor. Data type: int16 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N*block_shape[0]*block_shape[1], (H + pad_top + pad_bottom)/block_shape[0], (W + pad_left + pad_right)/block_shape[1], C_IN] |
output_offset | const int32_t | in | Zero offset for the output tensor. NOT USED. Assume symmetric quantization for s16. |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Batch to Space ND function for s8 data type.
arm_cmsis_nn_status arm_batch_to_space_nd_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const cmsis_nn_tile *block_shape, const cmsis_nn_dims *crop, int8_t *output_data, const cmsis_nn_dims *output_dims)Batch to Space ND function for s8 data type.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to the input tensor. Data type: int8 |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
block_shape | const cmsis_nn_tile * | in | Block shape for batch to space transformation |
crop | const cmsis_nn_dims * | in | Cropping for height and width. Format: [n->top, h->left, w->bottom, c->right] |
output_data | int8_t * | out | Pointer to the output tensor. Data type: int8 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N/(block_shape[0]*block_shape[1]), H*block_shape[0] - crop_top - crop_bottom, W*block_shape[1] - crop_left - crop_right, C_IN] |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Batch to Space ND function for s16 data type.
arm_cmsis_nn_status arm_batch_to_space_nd_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const cmsis_nn_tile *block_shape, const cmsis_nn_dims *crop, int16_t *output_data, const cmsis_nn_dims *output_dims)Batch to Space ND function for s16 data type.
- Supported Framework: TensorFlow Lite
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to the input tensor. Data type: int16 |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions. Format: [N, H, W, C_IN] |
block_shape | const cmsis_nn_tile * | in | Block shape for batch to space transformation |
crop | const cmsis_nn_dims * | in | Cropping for height and width. Format: [n->top, h->left, w->bottom, c->right] |
output_data | int16_t * | out | Pointer to the output tensor. Data type: int16 |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions. Format: [N/(block_shape[0]*block_shape[1]), H*block_shape[0] - crop_top - crop_bottom, W*block_shape[1] - crop_left - crop_right, C_IN] |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Basic transpose function.
arm_cmsis_nn_status arm_transpose_s8( const int8_t *input_data, int8_t *const output_data, const cmsis_nn_dims *const input_dims, const cmsis_nn_dims *const output_dims, const cmsis_nn_transpose_params *const transpose_params)Basic transpose function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Input (activation) data pointer. Data type: int8 |
output_data | int8_t *const | out | Output data pointer. Data type: int8 |
input_dims | const cmsis_nn_dims *const | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions. Format may be arbitrary relative to input format. The output dimension will depend on the permutation dimensions. In other words the out dimensions are the result of applying the permutation to the input dimensions. The first transpose_params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[permutations[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not. |
transpose_params | const cmsis_nn_transpose_params *const | in | Transpose parameters. Contains permutation dimensions. num_dims must be in [1, 4] and permutations must be a bijection over [0, num_dims - 1]. |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
Basic s16 transpose function.
arm_cmsis_nn_status arm_transpose_s16( const int16_t *input_data, int16_t *const output_data, const cmsis_nn_dims *const input_dims, const cmsis_nn_dims *const output_dims, const cmsis_nn_transpose_params *const transpose_params)Basic s16 transpose function.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Input (activation) data pointer. Data type: int16 |
output_data | int16_t *const | out | Output data pointer. Data type: int16 |
input_dims | const cmsis_nn_dims *const | in | Input (activation) tensor dimensions. Format: [N, H, W, C_IN] |
output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions. Format may be arbitrary relative to input format. The output dimension will depend on the permutation dimensions. In other words the out dimensions are the result of applying the permutation to the input dimensions. The first transpose_params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[permutations[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not. |
transpose_params | const cmsis_nn_transpose_params *const | in | Transpose parameters. Contains permutation dimensions. num_dims must be in [1, 4] and permutations must be a bijection over [0, num_dims - 1]. |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate.
void arm_concatenation_s8_x( const int8_t *input, const uint16_t input_x, const uint16_t input_y, const uint16_t input_z, const uint16_t input_w, int8_t *output, const uint16_t output_x, const uint32_t offset_x)int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate. The argument offset_x will be used to store the input tensor in the correct position in the output tensor
i.e. offset_x = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_x(&input[i], …, &output, …, …, offset_x) offset_x += input_x[i] }
This function assumes that the output tensor has:
- The same height of the input tensor
- The same number of channels of the input tensor
- The same batch size of the input tensor
Unless specified otherwise, arguments are mandatory.
Input constraints offset_x is less than output_x
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input tensor. Input tensor must not overlap with the output tensor. |
input_x | const uint16_t | in | Width of input tensor |
input_y | const uint16_t | in | Height of input tensor |
input_z | const uint16_t | in | Channels in input tensor |
input_w | const uint16_t | in | Batch size in input tensor |
output | int8_t * | out | Pointer to output tensor. Expected to be at least (input_x * input_y * input_z * input_w) + offset_x bytes. |
output_x | const uint16_t | in | Width of output tensor |
offset_x | const uint32_t | in | The offset (in number of elements) on the X axis to start concatenating the input tensor It is user responsibility to provide the correct value |
int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate.
void arm_concatenation_s8_y( const int8_t *input, const uint16_t input_x, const uint16_t input_y, const uint16_t input_z, const uint16_t input_w, int8_t *output, const uint16_t output_y, const uint32_t offset_y)int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate. The argument offset_y will be used to store the input tensor in the correct position in the output tensor
i.e. offset_y = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_y(&input[i], …, &output, …, …, offset_y) offset_y += input_y[i] }
This function assumes that the output tensor has:
- The same width of the input tensor
- The same number of channels of the input tensor
- The same batch size of the input tensor
Unless specified otherwise, arguments are mandatory.
Input constraints offset_y is less than output_y
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input tensor. Input tensor must not overlap with the output tensor. |
input_x | const uint16_t | in | Width of input tensor |
input_y | const uint16_t | in | Height of input tensor |
input_z | const uint16_t | in | Channels in input tensor |
input_w | const uint16_t | in | Batch size in input tensor |
output | int8_t * | out | Pointer to output tensor. Expected to be at least (input_z * input_w * input_x * input_y) + offset_y bytes. |
output_y | const uint16_t | in | Height of output tensor |
offset_y | const uint32_t | in | The offset on the Y axis to start concatenating the input tensor It is user responsibility to provide the correct value |
int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate.
void arm_concatenation_s8_z( const int8_t *input, const uint16_t input_x, const uint16_t input_y, const uint16_t input_z, const uint16_t input_w, int8_t *output, const uint16_t output_z, const uint32_t offset_z)int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate. The argument offset_z will be used to store the input tensor in the correct position in the output tensor
i.e. offset_z = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_z(&input[i], …, &output, …, …, offset_z) offset_z += input_z[i] }
This function assumes that the output tensor has:
- The same width of the input tensor
- The same height of the input tensor
- The same batch size of the input tensor
Unless specified otherwise, arguments are mandatory.
Input constraints offset_z is less than output_z
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input tensor. Input tensor must not overlap with output tensor. |
input_x | const uint16_t | in | Width of input tensor |
input_y | const uint16_t | in | Height of input tensor |
input_z | const uint16_t | in | Channels in input tensor |
input_w | const uint16_t | in | Batch size in input tensor |
output | int8_t * | out | Pointer to output tensor. Expected to be at least (input_x * input_y * input_z * input_w) + offset_z bytes. |
output_z | const uint16_t | in | Channels in output tensor |
offset_z | const uint32_t | in | The offset on the Z axis to start concatenating the input tensor It is user responsibility to provide the correct value |
int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to…
void arm_concatenation_s8_w( const int8_t *input, const uint16_t input_x, const uint16_t input_y, const uint16_t input_z, const uint16_t input_w, int8_t *output, const uint32_t offset_w)int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to concatenate. The argument offset_w will be used to store the input tensor in the correct position in the output tensor
i.e. offset_w = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_w(&input[i], …, &output, …, …, offset_w) offset_w += input_w[i] }
This function assumes that the output tensor has:
- The same width of the input tensor
- The same height of the input tensor
- The same number o channels of the input tensor
Unless specified otherwise, arguments are mandatory.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input tensor |
input_x | const uint16_t | in | Width of input tensor |
input_y | const uint16_t | in | Height of input tensor |
input_z | const uint16_t | in | Channels in input tensor |
input_w | const uint16_t | in | Batch size in input tensor |
output | int8_t * | out | Pointer to output tensor. Expected to be at least input_x * input_y * input_z * input_w bytes. |
offset_w | const uint32_t | in | The offset on the W axis to start concatenating the input tensor It is user responsibility to provide the correct value |
int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis
arm_cmsis_nn_status arm_concatenation_s8( const int8_t *const *input_data, const int32_t inputs_count, const int32_t *input_concat_dims, const int32_t axis, int8_t *output_data, const int32_t output_dims, const int32_t *output_shape)int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t *const * | in | Pointer to input tensors |
inputs_count | const int32_t | in | Number of input tensors |
input_concat_dims | const int32_t * | in | Dimensions of the input tensors along the target axis |
axis | const int32_t | in | Target axis to concatenate the input tensors |
output_data | int8_t * | out | Pointer to output tensor |
output_dims | const int32_t | in | Output tensor dimensions |
output_shape | const int32_t * | in | Output tensor shape |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis
arm_cmsis_nn_status arm_concatenation_s16( const int16_t *const *input_data, const int32_t inputs_count, const int32_t *input_concat_dims, const int32_t axis, int16_t *output_data, const int32_t output_dims, const int32_t *output_shape)int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t *const * | in | Pointer to input tensors |
inputs_count | const int32_t | in | Number of input tensors |
input_concat_dims | const int32_t * | in | Dimensions of the input tensors along the target axis |
axis | const int32_t | in | Target axis to concatenate the input tensors |
output_data | int16_t * | out | Pointer to output tensor |
output_dims | const int32_t | in | Output tensor dimensions |
output_shape | const int32_t * | in | Output tensor shape |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis
arm_cmsis_nn_status arm_concatenation_s32( const int32_t *const *input_data, const int32_t inputs_count, const int32_t *input_concat_dims, const int32_t axis, int32_t *output_data, const int32_t output_dims, const int32_t *output_shape)int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int32_t *const * | in | Pointer to input tensors |
inputs_count | const int32_t | in | Number of input tensors |
input_concat_dims | const int32_t * | in | Dimensions of the input tensors along the target axis |
axis | const int32_t | in | Target axis to concatenate the input tensors |
output_data | int32_t * | out | Pointer to output tensor |
output_dims | const int32_t | in | Output tensor dimensions |
output_shape | const int32_t * | in | Output tensor shape |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis
arm_cmsis_nn_status arm_split_s8( const int8_t *input_data, const int32_t input_dims, const int32_t *input_shape, const int32_t axis, const int32_t num_splits, const int32_t *split_dims, int8_t *const *output_data)int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to the flattened input tensor data. |
input_dims | const int32_t | in | Number of dimensions in input_shape. |
input_shape | const int32_t * | in | Array of length input_dims describing the shape of input_data. |
axis | const int32_t | in | Axis along which to split (0 <= axis < input_dims). |
num_splits | const int32_t | in | Number of output tensors to produce. |
split_dims | const int32_t * | in | Array of length num_splits giving size of each slice along axis. |
output_data | int8_t *const * | out | Array of pointers; output_data[i] points to storage for the i-th output tensor. |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR if split_dims sum mismatch. |
int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis
arm_cmsis_nn_status arm_split_s16( const int16_t *input_data, const int32_t input_dims, const int32_t *input_shape, const int32_t axis, const int32_t num_splits, const int32_t *split_dims, int16_t *const *output_data)int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to the flattened input tensor data. |
input_dims | const int32_t | in | Number of dimensions in input_shape. |
input_shape | const int32_t * | in | Array of length input_dims describing the shape of input_data. |
axis | const int32_t | in | Axis along which to split (0 <= axis < input_dims). |
num_splits | const int32_t | in | Number of output tensors to produce. |
split_dims | const int32_t * | in | Array of length num_splits giving size of each slice along axis. |
output_data | int16_t *const * | out | Array of pointers; output_data[i] points to storage for the i-th output tensor. |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR if split_dims sum mismatch. |
s8 SVDF function with 8 bit state tensor and 8 bit time weights
arm_cmsis_nn_status arm_svdf_s8( const cmsis_nn_context *ctx, const cmsis_nn_context *input_ctx, const cmsis_nn_context *output_ctx, const cmsis_nn_svdf_params *svdf_params, const cmsis_nn_per_tensor_quant_params *input_quant_params, const cmsis_nn_per_tensor_quant_params *output_quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *state_dims, int8_t *state_data, const cmsis_nn_dims *weights_feature_dims, const int8_t *weights_feature_data, const cmsis_nn_dims *weights_time_dims, const int8_t *weights_time_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 SVDF function with 8 bit state tensor and 8 bit time weights
- Supported framework: TensorFlow Lite micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Precomputed per-feature-batch kernel sums, supplied by the caller. This is an input the function only reads, not scratch it fills: an allocated but unfilled buffer yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. Mandatory on builds with the MVE extension (ARM_MATH_MVEI), where a NULL ctx->buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. Unused on every other build, where ctx->buf may be NULL. Sized by arm_svdf_s8_get_buffer_size(weights_feature_dims): weights_feature_dims->n * sizeof(int32_t) where the sums are used, 0 otherwise. Note this is weights_feature_dims->n, not a filter_dims->c - do not size this buffer with `arm_fully_connected_s8_get_buffer_size()`, which reads a different field and under-allocates. Fill it with arm_vector_sum_s8(ctx->buf, input_dims->h, weights_feature_dims->n, weights_feature_data, -svdf_params->input_offset, 0, NULL) so that entry j holds -input_offset * sum(weights_feature row j). The contents depend only on weights_feature_data and svdf_params->input_offset, so they may be computed once at load time and reused across calls until one of those changes. The buffer is specific to one layer's weights and cannot be shared between layers. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call. |
input_ctx | const cmsis_nn_context * | in | Scratch buffer written by this function, holding one int32_t accumulator per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_s8_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(int32_t) bytes, the same figure on every build target. This function does not read input_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons. |
output_ctx | const cmsis_nn_context * | in | Scratch buffer written by this function, holding one int32_t accumulator per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_s8_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t) bytes, truncating division, the same figure on every build target. This function does not read output_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons. |
svdf_params | const cmsis_nn_svdf_params * | in | SVDF Parameters Range of svdf_params->input_offset : [-128, 127] Range of svdf_params->output_offset : [-128, 127] |
input_quant_params | const cmsis_nn_per_tensor_quant_params * | in | Input quantization parameters |
output_quant_params | const cmsis_nn_per_tensor_quant_params * | in | Output quantization parameters |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
input_data | const int8_t * | in | Pointer to input tensor |
state_dims | const cmsis_nn_dims * | in | State tensor dimensions |
state_data | int8_t * | in, out | Pointer to state tensor |
weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions |
weights_feature_data | const int8_t * | in | Pointer to the weights (feature) tensor |
weights_time_dims | const cmsis_nn_dims * | in | Weights (time) tensor dimensions |
weights_time_data | const int8_t * | in | Pointer to the weights (time) tensor |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions |
bias_data | const int32_t * | in | Pointer to bias tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
output_data | int8_t * | out | Pointer to the output tensor |
Returns
| Description |
|---|
| The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion. |
s8 SVDF function with 16 bit state tensor and 16 bit time weights
arm_cmsis_nn_status arm_svdf_state_s16_s8( const cmsis_nn_context *input_ctx, const cmsis_nn_context *output_ctx, const cmsis_nn_svdf_params *svdf_params, const cmsis_nn_per_tensor_quant_params *input_quant_params, const cmsis_nn_per_tensor_quant_params *output_quant_params, const cmsis_nn_dims *input_dims, const int8_t *input_data, const cmsis_nn_dims *state_dims, int16_t *state_data, const cmsis_nn_dims *weights_feature_dims, const int8_t *weights_feature_data, const cmsis_nn_dims *weights_time_dims, const int16_t *weights_time_data, const cmsis_nn_dims *bias_dims, const int32_t *bias_data, const cmsis_nn_dims *output_dims, int8_t *output_data)s8 SVDF function with 16 bit state tensor and 16 bit time weights
- Supported framework: TensorFlow Lite micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_ctx | const cmsis_nn_context * | in | Scratch buffer written by this function, holding one int32_t accumulator per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_state_s16_s8_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(int32_t) bytes, the same figure on every build target. Note the accumulators are int32_t even though the state tensor is int16_t - this buffer does not shrink with the state width. This function does not read input_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons. |
output_ctx | const cmsis_nn_context * | in | Scratch buffer written by this function, holding one int32_t accumulator per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_state_s16_s8_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t) bytes, truncating division, the same figure on every build target. This function does not read output_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons. |
svdf_params | const cmsis_nn_svdf_params * | in | SVDF Parameters Range of svdf_params->input_offset : [-128, 127] Range of svdf_params->output_offset : [-128, 127] |
input_quant_params | const cmsis_nn_per_tensor_quant_params * | in | Input quantization parameters |
output_quant_params | const cmsis_nn_per_tensor_quant_params * | in | Output quantization parameters |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
input_data | const int8_t * | in | Pointer to input tensor |
state_dims | const cmsis_nn_dims * | in | State tensor dimensions |
state_data | int16_t * | in, out | Pointer to state tensor |
weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions |
weights_feature_data | const int8_t * | in | Pointer to the weights (feature) tensor |
weights_time_dims | const cmsis_nn_dims * | in | Weights (time) tensor dimensions |
weights_time_data | const int16_t * | in | Pointer to the weights (time) tensor |
bias_dims | const cmsis_nn_dims * | in | Bias tensor dimensions |
bias_data | const int32_t * | in | Pointer to bias tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
output_data | int8_t * | out | Pointer to the output tensor |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Get size of the kernel-sum buffer required by armsvdfs8().
int32_t arm_svdf_s8_get_buffer_size(const cmsis_nn_dims *weights_feature_dims)Get size of the kernel-sum buffer required by arm_svdf_s8().
For a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. arm_svdf_s8() has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other cmsis_nn_dims of that call can size it - in particular arm_fully_connected_s8_get_buffer_size() reads a different field and under-allocates. See arm_svdf_s8() for the buffer’s layout, how to fill it and when it may be reused.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
weights_feature_dims | const cmsis_nn_dims * | in | dimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t |
Get size of the kernel-sum buffer required by armsvdfs8() for processors with DSP extension.
int32_t arm_svdf_s8_get_buffer_size_dsp(const cmsis_nn_dims *weights_feature_dims)Get size of the kernel-sum buffer required by arm_svdf_s8() for processors with DSP extension.
For a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. arm_svdf_s8() has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other cmsis_nn_dims of that call can size it - in particular arm_fully_connected_s8_get_buffer_size() reads a different field and under-allocates. See arm_svdf_s8() for the buffer’s layout, how to fill it and when it may be reused.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
weights_feature_dims | const cmsis_nn_dims * | in | dimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t |
Get size of the kernel-sum buffer required by armsvdfs8() for Arm(R) Helium Architecture case.
int32_t arm_svdf_s8_get_buffer_size_mve(const cmsis_nn_dims *weights_feature_dims)Get size of the kernel-sum buffer required by arm_svdf_s8() for Arm(R) Helium Architecture case.
For a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. arm_svdf_s8() has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other cmsis_nn_dims of that call can size it - in particular arm_fully_connected_s8_get_buffer_size() reads a different field and under-allocates. See arm_svdf_s8() for the buffer’s layout, how to fill it and when it may be reused.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
weights_feature_dims | const cmsis_nn_dims * | in | dimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t |
Get size of the inputctx staging buffer required by armsvdfs8().
int32_t arm_svdf_s8_input_ctx_get_buffer_size( const cmsis_nn_dims *input_dims, const cmsis_nn_dims *weights_feature_dims)Get size of the input_ctx staging buffer required by arm_svdf_s8().
Returns input_dims->n * weights_feature_dims->n * sizeof(int32_t). Unlike arm_svdf_s8_get_buffer_size(), this figure does not vary by build target: arm_svdf_s8() stages this buffer on every build, not only under MVE, so there is no _dsp / _mve pair to choose between and the validation runs on every target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t |
Get size of the outputctx staging buffer required by armsvdfs8().
int32_t arm_svdf_s8_output_ctx_get_buffer_size( const cmsis_nn_svdf_params *svdf_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *weights_feature_dims)Get size of the output_ctx staging buffer required by arm_svdf_s8().
Returns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t). The division truncates, matching the kernel’s own unit count. As with arm_svdf_s8_input_ctx_get_buffer_size(), the figure is the same on every build target and the validation runs on every target.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
svdf_params | const cmsis_nn_svdf_params * | in | SVDF parameters; only svdf_params->rank is read |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any pointer is NULL, if svdf_params->rank is zero, negative or outside int16_t range, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t |
Get size of the inputctx staging buffer required by armsvdfstates16s8().
int32_t arm_svdf_state_s16_s8_input_ctx_get_buffer_size( const cmsis_nn_dims *input_dims, const cmsis_nn_dims *weights_feature_dims)Get size of the input_ctx staging buffer required by arm_svdf_state_s16_s8().
Returns input_dims->n * weights_feature_dims->n * sizeof(int32_t). Unlike arm_svdf_s8_get_buffer_size(), this figure does not vary by build target: arm_svdf_s8() stages this buffer on every build, not only under MVE, so there is no _dsp / _mve pair to choose between and the validation runs on every target.
Returns input_dims->n * weights_feature_dims->n * sizeof(int32_t) - the same figure as arm_svdf_s8_input_ctx_get_buffer_size() for the same shape. The accumulators are int32_t even though arm_svdf_state_s16_s8() carries an int16_t state tensor, so this buffer does not shrink with the state width.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t |
Get size of the outputctx staging buffer required by armsvdfstates16s8().
int32_t arm_svdf_state_s16_s8_output_ctx_get_buffer_size( const cmsis_nn_svdf_params *svdf_params, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *weights_feature_dims)Get size of the output_ctx staging buffer required by arm_svdf_state_s16_s8().
Returns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t). The division truncates, matching the kernel’s own unit count. As with arm_svdf_s8_input_ctx_get_buffer_size(), the figure is the same on every build target and the validation runs on every target.
Returns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t), truncating division - the same figure as arm_svdf_s8_output_ctx_get_buffer_size() for the same shape.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
svdf_params | const cmsis_nn_svdf_params * | in | SVDF parameters; only svdf_params->rank is read |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
weights_feature_dims | const cmsis_nn_dims * | in | Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if any pointer is NULL, if svdf_params->rank is zero, negative or outside int16_t range, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t |
LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias.
arm_cmsis_nn_status arm_lstm_unidirectional_s8( const int8_t *input, int8_t *output, const cmsis_nn_lstm_params *params, cmsis_nn_lstm_context *buffers)LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input data |
output | int8_t * | out | Pointer to output data |
params | const cmsis_nn_lstm_params * | in | Struct containing all information about the lstm operator, see arm_nn_types. |
buffers | cmsis_nn_lstm_context * | in, out | Struct containing pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. Size temp1 with `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` and temp2 with `arm_lstm_unidirectional_s8_temp2_get_buffer_size()` - both hold int16_t gate vectors even though the layer datatype is s8, so sizing them in s8 elements under-allocates by half. |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias.
arm_cmsis_nn_status arm_lstm_unidirectional_s16( const int16_t *input, int16_t *output, const cmsis_nn_lstm_params *params, cmsis_nn_lstm_context *buffers)LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to input data |
output | int16_t * | out | Pointer to output data |
params | const cmsis_nn_lstm_params * | in | Struct containing all information about the lstm operator, see arm_nn_types. |
buffers | cmsis_nn_lstm_context * | in, out | Struct containing pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. Size temp1 with `arm_lstm_unidirectional_s16_temp1_get_buffer_size()` and temp2 with `arm_lstm_unidirectional_s16_temp2_get_buffer_size()`. |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Get size of the temp1 scratch buffer required by armlstmunidirectionals8().
int32_t arm_lstm_unidirectional_s8_temp1_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)Get size of the temp1 scratch buffer required by arm_lstm_unidirectional_s8().
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
lstm_params | const cmsis_nn_lstm_params * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read. |
Returns
| Description |
|---|
| Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target. |
Get size of the temp2 scratch buffer required by armlstmunidirectionals8().
int32_t arm_lstm_unidirectional_s8_temp2_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)Get size of the temp2 scratch buffer required by arm_lstm_unidirectional_s8().
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
lstm_params | const cmsis_nn_lstm_params * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read. |
Returns
| Description |
|---|
| Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target. |
| Required buffer size in bytes: the same figure as `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` for the same params. temp2 stages the cell-gate vector and the tanh(cell_state) vector, both of the same extent as the gate vectors staged in temp1. |
Get size of the temp1 scratch buffer required by armlstmunidirectionals16().
int32_t arm_lstm_unidirectional_s16_temp1_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)Get size of the temp1 scratch buffer required by arm_lstm_unidirectional_s16().
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
lstm_params | const cmsis_nn_lstm_params * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read. |
Returns
| Description |
|---|
| Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target. |
| Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t) - the same figure as `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` for the same params, since both layer datatypes stage int16_t gate vectors. |
Get size of the temp2 scratch buffer required by armlstmunidirectionals16().
int32_t arm_lstm_unidirectional_s16_temp2_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)Get size of the temp2 scratch buffer required by arm_lstm_unidirectional_s16().
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
lstm_params | const cmsis_nn_lstm_params * | in | LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read. |
Returns
| Description |
|---|
| Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target. |
| Required buffer size in bytes: the same figure as `arm_lstm_unidirectional_s16_temp1_get_buffer_size()` for the same params. |
Batch matmul function with 8 bit input and output.
arm_cmsis_nn_status arm_batch_matmul_s8( const cmsis_nn_context *ctx, const cmsis_nn_bmm_params *bmm_params, const cmsis_nn_per_tensor_quant_params *quant_params, const cmsis_nn_dims *input_lhs_dims, const int8_t *input_lhs, const cmsis_nn_dims *input_rhs_dims, const int8_t *input_rhs, const cmsis_nn_dims *output_dims, int8_t *output)Batch matmul function with 8 bit input and output.
- Supported framework: TensorFlow Lite Micro
- Performs row * row matrix multiplication with the RHS transposed.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in, out | Temporary scratch buffer for the per-row kernel sums of the RHS. Mandatory on builds with the MVE extension (ARM_MATH_MVEI), where a NULL ctx->buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. Unused on every other build, where ctx->buf may be NULL. Sized by arm_batch_matmul_s8_get_buffer_size(input_rhs_dims) - pass the same input_rhs_dims given below. That is input_rhs_dims->w * sizeof(int32_t) where the sums are used, 0 otherwise. Do not size this buffer with `arm_fully_connected_s8_get_buffer_size()`: it reads a different field, and an allocation short of input_rhs_dims->w words is written past its end. The function fills the buffer itself before each use, so the caller does not need to initialize it. ctx->buf must be aligned to sizeof(int32_t). If ctx->size is non-zero it is validated against the requirement and a buffer too small is rejected with ARM_CMSIS_NN_ARG_ERROR; a ctx->size of 0 skips that check. A negative input_rhs_dims->w, or one large enough that the required size exceeds INT32_MAX, is rejected with ARM_CMSIS_NN_ARG_ERROR regardless of ctx->size. The caller is expected to clear the buffer, if applicable, for security reasons. |
bmm_params | const cmsis_nn_bmm_params * | in | Batch matmul Parameters Adjoint flags are currently unused and do not transpose either input; callers must supply the tensors in the layouts described below. |
quant_params | const cmsis_nn_per_tensor_quant_params * | in | Quantization parameters |
input_lhs_dims | const cmsis_nn_dims * | in | Input lhs tensor dimensions. This s8 function treats w as the row count and c as the inner dimension. This differs from `arm_batch_matmul_f32()`, so its dimension mapping must not be reused here. |
input_lhs | const int8_t * | in | Pointer to input tensor |
input_rhs_dims | const cmsis_nn_dims * | in | Input rhs tensor dimensions. The RHS must already be transposed, with w as its row count and c equal to input_lhs_dims->c. |
input_rhs | const int8_t * | in | Pointer to transposed input tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
output | int8_t * | out | Pointer to the output tensor |
Returns
| Description |
|---|
| The function returns one of the following: - `ARM_CMSIS_NN_ARG_ERROR` if an MVE build receives an invalid context, a negative or unrepresentable RHS row count, or a declared context size below the requirement. - `ARM_CMSIS_NN_SUCCESS` on success. |
Batch matmul function with 16 bit input and output.
arm_cmsis_nn_status arm_batch_matmul_s16( const cmsis_nn_context *ctx, const cmsis_nn_bmm_params *bmm_params, const cmsis_nn_per_tensor_quant_params *quant_params, const cmsis_nn_dims *input_lhs_dims, const int16_t *input_lhs, const cmsis_nn_dims *input_rhs_dims, const int16_t *input_rhs, const cmsis_nn_dims *output_dims, int16_t *output)Batch matmul function with 16 bit input and output.
- Supported framework: TensorFlow Lite Micro
- Performs row * row matrix multiplication with the RHS transposed.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
ctx | const cmsis_nn_context * | in | Unused: this function requires no scratch buffer and does not read or write ctx on any build, so ctx->buf may be NULL. Retained for signature compatibility with `arm_batch_matmul_s8()`. There is deliberately no arm_batch_matmul_s16_get_buffer_size(); in particular `arm_fully_connected_s8_get_buffer_size()` is not the sizer for this argument. If a real buffer is passed, the caller is expected to clear it, if applicable, for security reasons. |
bmm_params | const cmsis_nn_bmm_params * | in | Batch matmul Parameters Adjoint flags are currently unused. |
quant_params | const cmsis_nn_per_tensor_quant_params * | in | Quantization parameters |
input_lhs_dims | const cmsis_nn_dims * | in | Input lhs tensor dimensions. This should be NHWC where LHS.C = RHS.C |
input_lhs | const int16_t * | in | Pointer to input tensor |
input_rhs_dims | const cmsis_nn_dims * | in | Input lhs tensor dimensions. This is expected to be transposed so should be NHWC where LHS.C = RHS.C |
input_rhs | const int16_t * | in | Pointer to transposed input tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
output | int16_t * | out | Pointer to the output tensor |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Get size of the scratch buffer required by armbatchmatmuls8().
int32_t arm_batch_matmul_s8_get_buffer_size(const cmsis_nn_dims *input_rhs_dims)Get size of the scratch buffer required by arm_batch_matmul_s8().
For a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with arm_fully_connected_s8_get_buffer_size(), which reads .c) writes past the allocation whenever the rhs has more rows than columns. arm_batch_matmul_s16() needs no scratch buffer and so has no corresponding sizer.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_rhs_dims | const cmsis_nn_dims * | in | dimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t |
Get size of the scratch buffer required by armbatchmatmuls8() for processors with DSP extension.
int32_t arm_batch_matmul_s8_get_buffer_size_dsp(const cmsis_nn_dims *input_rhs_dims)Get size of the scratch buffer required by arm_batch_matmul_s8() for processors with DSP extension.
For a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with arm_fully_connected_s8_get_buffer_size(), which reads .c) writes past the allocation whenever the rhs has more rows than columns. arm_batch_matmul_s16() needs no scratch buffer and so has no corresponding sizer.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_rhs_dims | const cmsis_nn_dims * | in | dimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t |
Get size of the scratch buffer required by armbatchmatmuls8() for Arm(R) Helium Architecture case.
int32_t arm_batch_matmul_s8_get_buffer_size_mve(const cmsis_nn_dims *input_rhs_dims)Get size of the scratch buffer required by arm_batch_matmul_s8() for Arm(R) Helium Architecture case.
For a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with arm_fully_connected_s8_get_buffer_size(), which reads .c) writes past the allocation whenever the rhs has more rows than columns. arm_batch_matmul_s16() needs no scratch buffer and so has no corresponding sizer.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_rhs_dims | const cmsis_nn_dims * | in | dimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()` |
Returns
| Description |
|---|
| The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t |
Expands the size of the input by adding constant values before and after the data, in all dimensions.
arm_cmsis_nn_status arm_pad_s8( const int8_t *input, int8_t *output, const int8_t pad_value, const cmsis_nn_dims *input_size, const cmsis_nn_dims *pre_pad, const cmsis_nn_dims *post_pad)Expands the size of the input by adding constant values before and after the data, in all dimensions.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input data |
output | int8_t * | out | Pointer to output data |
pad_value | const int8_t | in | Value to pad with |
input_size | const cmsis_nn_dims * | in | Input tensor dimensions |
pre_pad | const cmsis_nn_dims * | in | Padding to apply before data in each dimension |
post_pad | const cmsis_nn_dims * | in | Padding to apply after data in each dimension |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Expands the size of the input by adding constant values before and after the data, in all dimensions.
arm_cmsis_nn_status arm_pad_s16( const int16_t *input, int16_t *output, const int16_t pad_value, const cmsis_nn_dims *input_size, const cmsis_nn_dims *pre_pad, const cmsis_nn_dims *post_pad)Expands the size of the input by adding constant values before and after the data, in all dimensions.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to input data |
output | int16_t * | out | Pointer to output data |
pad_value | const int16_t | in | Value to pad with |
input_size | const cmsis_nn_dims * | in | Input tensor dimensions |
pre_pad | const cmsis_nn_dims * | in | Padding to apply before data in each dimension |
post_pad | const cmsis_nn_dims * | in | Padding to apply after data in each dimension |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Computes the mean of the input tensor along the specified axis.
arm_cmsis_nn_status arm_mean_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const int32_t input_offset, const cmsis_nn_dims *axis_dims, int8_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift)Computes the mean of the input tensor along the specified axis. The output multipler and shift must have the output count folded into them.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
input_offset | const int32_t | in | Input offset |
axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
output_data | int8_t * | out | Pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
out_offset | const int32_t | in | Output offset |
out_mult | const int32_t | in | Output quantization multiplier |
out_shift | const int32_t | in | Output quantization shift |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Computes the mean of the input tensor along the specified axis.
arm_cmsis_nn_status arm_mean_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const int32_t input_offset, const cmsis_nn_dims *axis_dims, int16_t *output_data, const cmsis_nn_dims *output_dims, const int32_t out_offset, const int32_t out_mult, const int32_t out_shift)Computes the mean of the input tensor along the specified axis. The output multipler and shift must have the output count folded into them.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
input_offset | const int32_t | in | Input offset |
axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
output_data | int16_t * | out | Pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
out_offset | const int32_t | in | Output offset |
out_mult | const int32_t | in | Output quantization multiplier |
out_shift | const int32_t | in | Output quantization shift |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Compute ArgMax indices of an s8 tensor along a specific axis.
arm_cmsis_nn_status arm_argmax_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const int32_t axis, int32_t *output_data)Compute ArgMax indices of an s8 tensor along a specific axis.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (NHWC layout) |
axis | const int32_t | in | Reduction axis in range [0, 3] |
output_data | int32_t * | out | Pointer to output indices (int32_t) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`. |
Compute ArgMin indices of an s8 tensor along a specific axis.
arm_cmsis_nn_status arm_argmin_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const int32_t axis, int32_t *output_data)Compute ArgMin indices of an s8 tensor along a specific axis.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (NHWC layout) |
axis | const int32_t | in | Reduction axis in range [0, 3] |
output_data | int32_t * | out | Pointer to output indices (int32_t) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`. |
Compute ArgMax indices of an s16 tensor along a specific axis.
arm_cmsis_nn_status arm_argmax_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const int32_t axis, int32_t *output_data)Compute ArgMax indices of an s16 tensor along a specific axis.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (NHWC layout) |
axis | const int32_t | in | Reduction axis in range [0, 3] |
output_data | int32_t * | out | Pointer to output indices (int32_t) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`. |
Compute ArgMin indices of an s16 tensor along a specific axis.
arm_cmsis_nn_status arm_argmin_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const int32_t axis, int32_t *output_data)Compute ArgMin indices of an s16 tensor along a specific axis.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions (NHWC layout) |
axis | const int32_t | in | Reduction axis in range [0, 3] |
output_data | int32_t * | out | Pointer to output indices (int32_t) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`. |
Computes the max of the input tensor along the specified axis.
arm_cmsis_nn_status arm_reduce_max_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *axis_dims, int8_t *output_data, const cmsis_nn_dims *output_dims)Computes the max of the input tensor along the specified axis.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
output_data | int8_t * | out | Pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Computes the max of the input tensor along the specified axis.
arm_cmsis_nn_status arm_reduce_max_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *axis_dims, int16_t *output_data, const cmsis_nn_dims *output_dims)Computes the max of the input tensor along the specified axis.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
output_data | int16_t * | out | Pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Computes the min of the input tensor along the specified axis.
arm_cmsis_nn_status arm_reduce_min_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *axis_dims, int8_t *output_data, const cmsis_nn_dims *output_dims)Computes the min of the input tensor along the specified axis.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
output_data | int8_t * | out | Pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Computes the min of the input tensor along the specified axis.
arm_cmsis_nn_status arm_reduce_min_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const cmsis_nn_dims *axis_dims, int16_t *output_data, const cmsis_nn_dims *output_dims)Computes the min of the input tensor along the specified axis.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to input tensor |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
axis_dims | const cmsis_nn_dims * | in | Axis dimensions to compute mean over |
output_data | int16_t * | out | Pointer to output tensor |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Quantize a floating-point array into int8t format.
arm_cmsis_nn_status arm_quantize_f32_s8( const float *input, int8_t *output, int32_t size, int32_t zero_point, float scale)Quantize a floating-point array into int8_t format.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const float * | in | Pointer to the input float array. |
output | int8_t * | out | Pointer to the output int8_t array. |
size | int32_t | in | Number of elements in the arrays. |
zero_point | int32_t | in | Zero point (offset) to apply during quantization. |
scale | float | in | Scale factor to apply during quantization. Must be a positive finite number. A scale that is zero, negative, NaN, Inf, or small enough that its reciprocal overflows is unsupported, and the result is then unspecified: the scalar and Helium legs are not guaranteed to agree for such a scale. Denormal inputs are likewise unspecified: the Helium leg flushes them to zero, so the two legs are not guaranteed to agree for a denormal input at any scale. Screen denormals if you need them to match. |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when `zero_point` lies outside the int8_t range. Values round half away from zero and saturate to the int8_t range after the zero point is applied; NaN maps to `zero_point`. |
Quantize a floating-point array into int16t format.
arm_cmsis_nn_status arm_quantize_f32_s16( const float *input, int16_t *output, int32_t size, int32_t zero_point, float scale)Quantize a floating-point array into int16_t format.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const float * | in | Pointer to the input float array. |
output | int16_t * | out | Pointer to the output int16_t array. |
size | int32_t | in | Number of elements in the arrays. |
zero_point | int32_t | in | Zero point (offset) to apply during quantization. |
scale | float | in | Scale factor to apply during quantization. Must be a positive finite number. A scale that is zero, negative, NaN, Inf, or small enough that its reciprocal overflows is unsupported, and the result is then unspecified: the scalar and Helium legs are not guaranteed to agree for such a scale. Denormal inputs are likewise unspecified: the Helium leg flushes them to zero, so the two legs are not guaranteed to agree for a denormal input at any scale. Screen denormals if you need them to match. |
Returns
| Description |
|---|
| ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when `zero_point` lies outside the int16_t range. Values round half away from zero and saturate to the int16_t range after the zero point is applied; NaN maps to `zero_point`. |
Requantize an int8t array to another int8t range with a different scale.
arm_cmsis_nn_status arm_requantize_s8_s8( const int8_t *input, int8_t *output, int32_t size, int32_t effective_scale_multiplier, int32_t effective_scale_shift, int32_t input_zeropoint, int32_t output_zeropoint)Requantize an int8_t array to another int8_t range with a different scale.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input int8_t array. |
output | int8_t * | out | Pointer to the output int8_t array. |
size | int32_t | in | Number of elements in the arrays. |
effective_scale_multiplier | int32_t | in | Multiplier used for the scaling operation. |
effective_scale_shift | int32_t | in | Right or left shift (depending on sign) applied after the multiplier. |
input_zeropoint | int32_t | in | Zero point of the input data. |
output_zeropoint | int32_t | in | Zero point of the output data. |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Requantize an int16t array to another int16t range with a different scale.
arm_cmsis_nn_status arm_requantize_s16_s16( const int16_t *input, int16_t *output, int32_t size, int32_t effective_scale_multiplier, int32_t effective_scale_shift, int32_t input_zeropoint, int32_t output_zeropoint)Requantize an int16_t array to another int16_t range with a different scale.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input int16_t array. |
output | int16_t * | out | Pointer to the output int16_t array. |
size | int32_t | in | Number of elements in the arrays. |
effective_scale_multiplier | int32_t | in | Multiplier used for the scaling operation. |
effective_scale_shift | int32_t | in | Right or left shift (depending on sign) applied after the multiplier. |
input_zeropoint | int32_t | in | Zero point of the input data. |
output_zeropoint | int32_t | in | Zero point of the output data. |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Dequantize an int8t array back to floating-point format.
arm_cmsis_nn_status arm_dequantize_s8_f32( const int8_t *input, float *output, int32_t size, int32_t zero_point, float scale)Dequantize an int8_t array back to floating-point format.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to the input int8_t array. |
output | float * | out | Pointer to the output float array. |
size | int32_t | in | Number of elements in the arrays. |
zero_point | int32_t | in | Zero point (offset) that was used during quantization. |
scale | float | in | Scale factor that was used during quantization. |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Dequantize an int16t array back to floating-point format.
arm_cmsis_nn_status arm_dequantize_s16_f32( const int16_t *input, float *output, int32_t size, int32_t zero_point, float scale)Dequantize an int16_t array back to floating-point format.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to the input int16_t array. |
output | float * | out | Pointer to the output float array. |
size | int32_t | in | Number of elements in the arrays. |
zero_point | int32_t | in | Zero point (offset) that was used during quantization. |
scale | float | in | Scale factor that was used during quantization. |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Strided slice function for int8 data.
arm_cmsis_nn_status arm_strided_slice_s8( const int8_t *input_data, int8_t *output_data, const cmsis_nn_dims *const input_dims, const cmsis_nn_dims *const begin_dims, const cmsis_nn_dims *const stride_dims, const cmsis_nn_dims *const output_dims)Strided slice function for int8 data.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to input tensor |
output_data | int8_t * | out | Pointer to output tensor |
input_dims | const cmsis_nn_dims *const | in | Input tensor dimensions |
begin_dims | const cmsis_nn_dims *const | in | Begin dimensions for slicing |
stride_dims | const cmsis_nn_dims *const | in | Stride dimensions for slicing |
output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Strided slice function for int16 data.
arm_cmsis_nn_status arm_strided_slice_s16( const int16_t *input_data, int16_t *output_data, const cmsis_nn_dims *const input_dims, const cmsis_nn_dims *const begin_dims, const cmsis_nn_dims *const stride_dims, const cmsis_nn_dims *const output_dims)Strided slice function for int16 data.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to input tensor |
output_data | int16_t * | out | Pointer to output tensor |
input_dims | const cmsis_nn_dims *const | in | Input tensor dimensions |
begin_dims | const cmsis_nn_dims *const | in | Begin dimensions for slicing |
stride_dims | const cmsis_nn_dims *const | in | Stride dimensions for slicing |
output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Strided slice function for int32 data.
arm_cmsis_nn_status arm_strided_slice_s32( const int32_t *input_data, int32_t *output_data, const cmsis_nn_dims *const input_dims, const cmsis_nn_dims *const begin_dims, const cmsis_nn_dims *const stride_dims, const cmsis_nn_dims *const output_dims)Strided slice function for int32 data.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int32_t * | in | Pointer to input tensor |
output_data | int32_t * | out | Pointer to output tensor |
input_dims | const cmsis_nn_dims *const | in | Input tensor dimensions |
begin_dims | const cmsis_nn_dims *const | in | Begin dimensions for slicing |
stride_dims | const cmsis_nn_dims *const | in | Stride dimensions for slicing |
output_dims | const cmsis_nn_dims *const | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Gather elements along an axis for int8 tensors.
arm_cmsis_nn_status arm_gather_s8( const int8_t *input_data, const cmsis_nn_dims *input_dims, const int32_t *indices_data, const cmsis_nn_dims *indices_dims, const cmsis_nn_gather_params *params, int8_t *output_data, const cmsis_nn_dims *output_dims)Gather elements along an axis for int8 tensors.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int8_t * | in | Pointer to input tensor data |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
indices_data | const int32_t * | in | Pointer to indices tensor data (int32) |
indices_dims | const cmsis_nn_dims * | in | Indices tensor dimensions |
params | const cmsis_nn_gather_params * | in | Pointer to gather parameters |
output_data | int8_t * | out | Pointer to output tensor data |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Gather elements along an axis for int16 tensors.
arm_cmsis_nn_status arm_gather_s16( const int16_t *input_data, const cmsis_nn_dims *input_dims, const int32_t *indices_data, const cmsis_nn_dims *indices_dims, const cmsis_nn_gather_params *params, int16_t *output_data, const cmsis_nn_dims *output_dims)Gather elements along an axis for int16 tensors.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input_data | const int16_t * | in | Pointer to input tensor data |
input_dims | const cmsis_nn_dims * | in | Input tensor dimensions |
indices_data | const int32_t * | in | Pointer to indices tensor data (int32) |
indices_dims | const cmsis_nn_dims * | in | Indices tensor dimensions |
params | const cmsis_nn_gather_params * | in | Pointer to gather parameters |
output_data | int16_t * | out | Pointer to output tensor data |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Gathernd slices for int8 tensors.
arm_cmsis_nn_status arm_gather_nd_s8( const int8_t *params_data, const cmsis_nn_dims *params_dims, const int32_t *indices_data, const cmsis_nn_dims *indices_dims, const cmsis_nn_gather_nd_params *params, int8_t *output_data, const cmsis_nn_dims *output_dims)Gather_nd slices for int8 tensors.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
params_data | const int8_t * | in | Pointer to params tensor data |
params_dims | const cmsis_nn_dims * | in | Params tensor dimensions |
indices_data | const int32_t * | in | Pointer to indices tensor data (int32) |
indices_dims | const cmsis_nn_dims * | in | Indices tensor dimensions |
params | const cmsis_nn_gather_nd_params * | in | Pointer to gather_nd parameters |
output_data | int8_t * | out | Pointer to output tensor data |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Gathernd slices for int16 tensors.
arm_cmsis_nn_status arm_gather_nd_s16( const int16_t *params_data, const cmsis_nn_dims *params_dims, const int32_t *indices_data, const cmsis_nn_dims *indices_dims, const cmsis_nn_gather_nd_params *params, int16_t *output_data, const cmsis_nn_dims *output_dims)Gather_nd slices for int16 tensors.
- Supported framework: TensorFlow Lite Micro
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
params_data | const int16_t * | in | Pointer to params tensor data |
params_dims | const cmsis_nn_dims * | in | Params tensor dimensions |
indices_data | const int32_t * | in | Pointer to indices tensor data (int32) |
indices_dims | const cmsis_nn_dims * | in | Indices tensor dimensions |
params | const cmsis_nn_gather_nd_params * | in | Pointer to gather_nd parameters |
output_data | int16_t * | out | Pointer to output tensor data |
output_dims | const cmsis_nn_dims * | in | Output tensor dimensions |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Tile an int8 tensor along each dimension.
arm_cmsis_nn_status arm_tile_s8(const int8_t *input, const cmsis_nn_tile_params *params, int8_t *output)Tile an int8 tensor along each dimension.
- Supported framework: TensorFlow Lite Micro
- Maximum rank: 8
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input tensor data |
params | const cmsis_nn_tile_params * | in | Pointer to tile parameters (rank, input_shape, multiples) |
output | int8_t * | out | Pointer to output tensor data (pre-allocated by caller) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Tile an int16 tensor along each dimension.
arm_cmsis_nn_status arm_tile_s16(const int16_t *input, const cmsis_nn_tile_params *params, int16_t *output)Tile an int16 tensor along each dimension.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to input tensor data |
params | const cmsis_nn_tile_params * | in | Pointer to tile parameters (rank, input_shape, multiples) |
output | int16_t * | out | Pointer to output tensor data (pre-allocated by caller) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Broadcast an int8 tensor to a target shape.
arm_cmsis_nn_status arm_broadcast_to_s8(const int8_t *input, const cmsis_nn_broadcast_to_params *params, int8_t *output)Broadcast an int8 tensor to a target shape.
- Input dimensions must be 1 or match the output dimension for each axis.
- Maximum rank: 8
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input tensor data |
params | const cmsis_nn_broadcast_to_params * | in | Pointer to broadcast parameters (rank, input/output shapes) |
output | int8_t * | out | Pointer to output tensor data (pre-allocated by caller) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Broadcast an int16 tensor to a target shape.
arm_cmsis_nn_status arm_broadcast_to_s16( const int16_t *input, const cmsis_nn_broadcast_to_params *params, int16_t *output)Broadcast an int16 tensor to a target shape.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to input tensor data |
params | const cmsis_nn_broadcast_to_params * | in | Pointer to broadcast parameters (rank, input/output shapes) |
output | int16_t * | out | Pointer to output tensor data (pre-allocated by caller) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Scatter updates into a zero-initialized output tensor for int8.
arm_cmsis_nn_status arm_scatter_nd_s8( const int32_t *indices, const int8_t *updates, const cmsis_nn_scatter_nd_params *params, int8_t *output)Scatter updates into a zero-initialized output tensor for int8.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
indices | const int32_t * | in | Pointer to indices data (int32, shape [num_updates, index_depth]) |
updates | const int8_t * | in | Pointer to updates data |
params | const cmsis_nn_scatter_nd_params * | in | Pointer to scatter_nd parameters |
output | int8_t * | out | Pointer to output tensor data (pre-allocated by caller) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Scatter updates into a zero-initialized output tensor for int16.
arm_cmsis_nn_status arm_scatter_nd_s16( const int32_t *indices, const int16_t *updates, const cmsis_nn_scatter_nd_params *params, int16_t *output)Scatter updates into a zero-initialized output tensor for int16.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
indices | const int32_t * | in | Pointer to indices data (int32, shape [num_updates, index_depth]) |
updates | const int16_t * | in | Pointer to updates data |
params | const cmsis_nn_scatter_nd_params * | in | Pointer to scatter_nd parameters |
output | int16_t * | out | Pointer to output tensor data (pre-allocated by caller) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode).
arm_cmsis_nn_status arm_mirror_pad_s8(const int8_t *input, const cmsis_nn_mirror_pad_params *params, int8_t *output)Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode).
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input tensor data |
params | const cmsis_nn_mirror_pad_params * | in | Pointer to mirror_pad parameters |
output | int8_t * | out | Pointer to output tensor data (pre-allocated by caller) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode).
arm_cmsis_nn_status arm_mirror_pad_s16(const int16_t *input, const cmsis_nn_mirror_pad_params *params, int16_t *output)Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode).
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to input tensor data |
params | const cmsis_nn_mirror_pad_params * | in | Pointer to mirror_pad parameters |
output | int16_t * | out | Pointer to output tensor data (pre-allocated by caller) |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
WHERE operator: return coordinates of non-zero elements in condition.
arm_cmsis_nn_status arm_where_s8( const int8_t *condition, const cmsis_nn_where_params *params, int64_t *output, int32_t *num_true)WHERE operator: return coordinates of non-zero elements in condition.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
condition | const int8_t * | in | Pointer to condition tensor data (int8, non-zero = true) |
params | const cmsis_nn_where_params * | in | Pointer to where parameters (rank, shape) |
output | int64_t * | out | Pointer to output coordinates (int64, shape [max_true, rank]) |
num_true | int32_t * | out | Number of true elements found |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
WHERE operator: return coordinates of non-zero elements in condition (int16).
arm_cmsis_nn_status arm_where_s16( const int16_t *condition, const cmsis_nn_where_params *params, int64_t *output, int32_t *num_true)WHERE operator: return coordinates of non-zero elements in condition (int16).
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
condition | const int16_t * | in | Pointer to condition tensor data (int16, non-zero = true) |
params | const cmsis_nn_where_params * | in | Pointer to where parameters (rank, shape) |
output | int64_t * | out | Pointer to output coordinates (int64, shape [max_true, rank]) |
num_true | int32_t * | out | Number of true elements found |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
SELECTV2 with broadcast for int8 tensors.
arm_cmsis_nn_status arm_select_v2_s8( const bool *condition, const int8_t *x, const int8_t *y, const cmsis_nn_select_v2_params *params, int8_t *output)SELECT_V2 with broadcast for int8 tensors.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
condition | const bool * | in | Pointer to condition tensor data (bool) |
x | const int8_t * | in | Pointer to x tensor data (selected when condition is true) |
y | const int8_t * | in | Pointer to y tensor data (selected when condition is false) |
params | const cmsis_nn_select_v2_params * | in | Pointer to select_v2 parameters (broadcast strides) |
output | int8_t * | out | Pointer to output tensor data |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
SELECTV2 with broadcast for int16 tensors.
arm_cmsis_nn_status arm_select_v2_s16( const bool *condition, const int16_t *x, const int16_t *y, const cmsis_nn_select_v2_params *params, int16_t *output)SELECT_V2 with broadcast for int16 tensors.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
condition | const bool * | in | Pointer to condition tensor data (bool) |
x | const int16_t * | in | Pointer to x tensor data |
y | const int16_t * | in | Pointer to y tensor data |
params | const cmsis_nn_select_v2_params * | in | Pointer to select_v2 parameters (broadcast strides) |
output | int16_t * | out | Pointer to output tensor data |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Reverse variable-length sequences along a dimension for int8.
arm_cmsis_nn_status arm_reverse_sequence_s8( const int8_t *input, const int32_t *seq_lengths, const cmsis_nn_reverse_sequence_params *params, int8_t *output)Reverse variable-length sequences along a dimension for int8.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int8_t * | in | Pointer to input tensor data |
seq_lengths | const int32_t * | in | Pointer to per-batch sequence lengths (int32) |
params | const cmsis_nn_reverse_sequence_params * | in | Pointer to reverse_sequence parameters |
output | int8_t * | out | Pointer to output tensor data |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Reverse variable-length sequences along a dimension for int16.
arm_cmsis_nn_status arm_reverse_sequence_s16( const int16_t *input, const int32_t *seq_lengths, const cmsis_nn_reverse_sequence_params *params, int16_t *output)Reverse variable-length sequences along a dimension for int16.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
input | const int16_t * | in | Pointer to input tensor data |
seq_lengths | const int32_t * | in | Pointer to per-batch sequence lengths (int32) |
params | const cmsis_nn_reverse_sequence_params * | in | Pointer to reverse_sequence parameters |
output | int16_t * | out | Pointer to output tensor data |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Update a slice of an int8 operand tensor at runtime-determined indices.
arm_cmsis_nn_status arm_dynamic_update_slice_s8( const int8_t *operand, const int8_t *update, const int32_t *start_indices, const cmsis_nn_dynamic_update_slice_params *params, int8_t *output)Update a slice of an int8 operand tensor at runtime-determined indices.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
operand | const int8_t * | in | Pointer to operand tensor data (copied to output first) |
update | const int8_t * | in | Pointer to update tensor data |
start_indices | const int32_t * | in | Pointer to start index per dimension (int32, length = rank) |
params | const cmsis_nn_dynamic_update_slice_params * | in | Pointer to dynamic_update_slice parameters |
output | int8_t * | out | Pointer to output tensor data |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |
Update a slice of an int16 operand tensor at runtime-determined indices.
arm_cmsis_nn_status arm_dynamic_update_slice_s16( const int16_t *operand, const int16_t *update, const int32_t *start_indices, const cmsis_nn_dynamic_update_slice_params *params, int16_t *output)Update a slice of an int16 operand tensor at runtime-determined indices.
Parameters
| Name | Type | Direction | Description |
|---|---|---|---|
operand | const int16_t * | in | Pointer to operand tensor data (copied to output first) |
update | const int16_t * | in | Pointer to update tensor data |
start_indices | const int32_t * | in | Pointer to start index per dimension (int32, length = rank) |
params | const cmsis_nn_dynamic_update_slice_params * | in | Pointer to dynamic_update_slice parameters |
output | int16_t * | out | Pointer to output tensor data |
Returns
| Description |
|---|
| The function returns `ARM_CMSIS_NN_SUCCESS` |