Kernel index
Search by function name or description, then narrow the results by operator group and data type. Select a function to read its parameters and requirements.
426 of 426 functions
| Function | Description |
|---|---|
| arm_abs_s16Elementwise | s16 elementwise absolute value |
| arm_abs_s8Elementwise | s8 elementwise absolute value |
| arm_add_s16Elementwise | s16 elementwise add of two tensors with support for broadcasting. |
| arm_add_s8Elementwise | s8 elementwise add of two tensors with support for broadcasting. |
| arm_add_scalar_s16Elementwise | s16 elementwise add of scalar and vector |
| arm_add_scalar_s8Elementwise | s8 elementwise add of scalar and vector |
| arm_argmax_f16Reduction and comparison | Returns the first maximum's axis-relative INT32 index for a f16 tensor. |
| arm_argmax_f32Reduction and comparison | Returns the first maximum's axis-relative INT32 index for a f32 tensor. |
| arm_argmax_s16Reduction and comparison | Compute ArgMax indices of an s16 tensor along a specific axis. |
| arm_argmax_s8Reduction and comparison | Compute ArgMax indices of an s8 tensor along a specific axis. |
| arm_argmin_f16Reduction and comparison | Returns the first minimum's axis-relative INT32 index for a f16 tensor. |
| arm_argmin_f32Reduction and comparison | Returns the first minimum's axis-relative INT32 index for a f32 tensor. |
| arm_argmin_s16Reduction and comparison | Compute ArgMin indices of an s16 tensor along a specific axis. |
| arm_argmin_s8Reduction and comparison | Compute ArgMin indices of an s8 tensor along a specific axis. |
| arm_avg_pool_f16Pooling, softmax, quantization | Average pooling. |
| arm_avg_pool_f32Pooling, softmax, quantization | Average pooling. |
| arm_avgpool_s16Pooling, softmax, quantization | s16 average pooling function. |
| arm_avgpool_s16_get_buffer_sizePooling, softmax, quantization | Get the required buffer size for S16 average pooling function. |
| arm_avgpool_s16_get_buffer_size_dspPooling, softmax, quantization | Get the required buffer size for S16 average pooling function for processors with DSP extension. |
| arm_avgpool_s16_get_buffer_size_mvePooling, softmax, quantization | Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case. |
| arm_avgpool_s8Pooling, softmax, quantization | s8 average pooling function. |
| arm_avgpool_s8_get_buffer_sizePooling, softmax, quantization | Get the required buffer size for S8 average pooling function. |
| arm_avgpool_s8_get_buffer_size_dspPooling, softmax, quantization | Get the required buffer size for S8 average pooling function for processors with DSP extension. |
| arm_avgpool_s8_get_buffer_size_mvePooling, softmax, quantization | Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case. |
| arm_batch_matmul_f16Fully connected | Batched matrix multiplication. |
| arm_batch_matmul_f16_get_buffer_sizeFully connected | Get the temporary buffer size required by batched matrix multiplication. |
| arm_batch_matmul_f32Fully connected | Batched matrix multiplication. |
| arm_batch_matmul_f32_get_buffer_sizeFully connected | Get the temporary buffer size required by batched matrix multiplication. |
| arm_batch_matmul_s16Fully connected | Batch matmul function with 16 bit input and output. |
| arm_batch_matmul_s8Fully connected | Batch matmul function with 8 bit input and output. |
| arm_batch_matmul_s8_get_buffer_sizeFully connected | Get size of the scratch buffer required by armbatchmatmuls8(). |
| arm_batch_matmul_s8_get_buffer_size_dspFully connected | Get size of the scratch buffer required by armbatchmatmuls8() for processors with DSP extension. |
| arm_batch_matmul_s8_get_buffer_size_mveFully connected | Get size of the scratch buffer required by armbatchmatmuls8() for Arm(R) Helium Architecture case. |
| arm_batch_norm_f16Elementwise | Apply batch normalization. |
| arm_batch_norm_f32Elementwise | Apply batch normalization. |
| arm_batch_to_space_nd_s16Data movement | Batch to Space ND function for s16 data type. |
| arm_batch_to_space_nd_s8Data movement | Batch to Space ND function for s8 data type. |
| arm_broadcast_to_s16Data movement | Broadcast an int16 tensor to a target shape. |
| arm_broadcast_to_s8Data movement | Broadcast an int8 tensor to a target shape. |
| arm_clamp_s16Activation | S16 clamp function. |
| arm_clamp_s8Activation | S8 clamp function. |
| arm_comparison_s16Reduction and comparison | s16 elementwise comparison with support for broadcasting. |
| arm_comparison_s8Reduction and comparison | s8 elementwise comparison with support for broadcasting. |
| arm_concatenation_f16Data movement | Concatenate float32 tensors of any rank along one axis. |
| arm_concatenation_f16_wData movement | Concatenate tensors along the W axis. |
| arm_concatenation_f16_xData movement | Concatenate tensors along the X axis. |
| arm_concatenation_f16_yData movement | Concatenate tensors along the Y axis. |
| arm_concatenation_f16_zData movement | Concatenate tensors along the Z axis. |
| arm_concatenation_f32Data movement | Concatenate float32 tensors of any rank along one axis. |
| arm_concatenation_f32_wData movement | Concatenate tensors along the W axis. |
| arm_concatenation_f32_xData movement | Concatenate tensors along the X axis. |
| arm_concatenation_f32_yData movement | Concatenate tensors along the Y axis. |
| arm_concatenation_f32_zData movement | Concatenate tensors along the Z axis. |
| arm_concatenation_s16Data movement | int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis |
| arm_concatenation_s32Data movement | int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis |
| arm_concatenation_s8Data movement | int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis |
| arm_concatenation_s8_wData movement | int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to… |
| arm_concatenation_s8_xData movement | int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate. |
| arm_concatenation_s8_yData movement | int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate. |
| arm_concatenation_s8_zData movement | int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate. |
| arm_convolve_1_x_n_f16Convolution | 1xN convolution, dispatch by layout. |
| arm_convolve_1_x_n_f16_acc16Convolution | 1xN convolution, dispatch by layout. |
| arm_convolve_1_x_n_f16_get_buffer_sizeConvolution | Get the buffer size required by 1xN convolution. |
| arm_convolve_1_x_n_f32Convolution | 1xN convolution, dispatch by layout. |
| arm_convolve_1_x_n_f32_get_buffer_sizeConvolution | Get the buffer size required by 1xN convolution. |
| arm_convolve_1_x_n_nhwc_f16Convolution | 1xN convolution, NHWC layout. |
| arm_convolve_1_x_n_nhwc_f16_acc16Convolution | 1xN convolution, NHWC layout. |
| arm_convolve_1_x_n_nhwc_f32Convolution | 1xN convolution, NHWC layout. |
| arm_convolve_1_x_n_s4Convolution | 1xn convolution for s4 weights |
| arm_convolve_1_x_n_s4_get_buffer_sizeConvolution | Get the required additional buffer size for 1xn convolution. |
| arm_convolve_1_x_n_s8Convolution | 1xn convolution |
| arm_convolve_1_x_n_s8_get_buffer_sizeConvolution | Get the required additional buffer size for 1xn convolution. |
| arm_convolve_1x1_f16Convolution | 1x1 convolution, dispatch by layout. |
| arm_convolve_1x1_f16_acc16Convolution | 1x1 convolution, dispatch by layout. |
| arm_convolve_1x1_f16_get_buffer_sizeConvolution | Get the buffer size required by 1x1 convolution. |
| arm_convolve_1x1_f32Convolution | 1x1 convolution, dispatch by layout. |
| arm_convolve_1x1_f32_get_buffer_sizeConvolution | Get the buffer size required by 1x1 convolution. |
| arm_convolve_1x1_nhwc_f16Convolution | 1x1 convolution, NHWC layout. |
| arm_convolve_1x1_nhwc_f16_acc16Convolution | 1x1 convolution, NHWC layout. |
| arm_convolve_1x1_nhwc_f32Convolution | 1x1 convolution, NHWC layout. |
| arm_convolve_1x1_out_s8Convolution | Optimised convolution for 1x1 output images (shape of BX1x1xCOUT) for 8x8 computations. |
| arm_convolve_1x1_out_s8_get_buffer_sizeConvolution | Get the required scratch buffer size for armconvolve1x1outs8(). |
| arm_convolve_1x1_s16_ns_np_ndConvolution | Pointwise s16 convolution function: no stride, no padding, no dilation. |
| arm_convolve_1x1_s4Convolution | s4 version for 1x1 convolution with support for non-unity stride values |
| arm_convolve_1x1_s4_fastConvolution | Fast s4 version for 1x1 convolution (non-square shape). |
| arm_convolve_1x1_s4_fast_get_buffer_sizeConvolution | Get the required buffer size for armconvolve1x1s4fast. |
| arm_convolve_1x1_s8Convolution | s8 version for 1x1 convolution with support for non-unity stride values |
| arm_convolve_1x1_s8_fastConvolution | Fast s8 version for 1x1 convolution (non-square shape). |
| arm_convolve_1x1_s8_fast_get_buffer_sizeConvolution | Get the required buffer size for armconvolve1x1s8fast. |
| arm_convolve_even_s4Convolution | Basic s4 convolution function with a requirement of even number of kernels. |
| arm_convolve_even_s4_get_buffer_sizeConvolution | Get the required buffer size for armconvolveevens4. |
| arm_convolve_f16Convolution | Convolution, dispatch by layout. |
| arm_convolve_f16_acc16Convolution | Convolution, dispatch by layout. |
| arm_convolve_f16_get_buffer_sizeConvolution | Get the temporary buffer size required by convolution. |
| arm_convolve_f32Convolution | Convolution, dispatch by layout. |
| arm_convolve_f32_get_buffer_sizeConvolution | Get the temporary buffer size required by convolution. |
| arm_convolve_nhwc_f16Convolution | Convolution, NHWC layout. |
| arm_convolve_nhwc_f16_acc16Convolution | Convolution, NHWC layout. |
| arm_convolve_nhwc_f32Convolution | Convolution, NHWC layout. |
| arm_convolve_s16Convolution | Basic s16 convolution function. |
| arm_convolve_s16_fast_small_kernelConvolution | armconvolves16fastsmallkernel function. |
| arm_convolve_s16_get_buffer_sizeConvolution | Get the required buffer size for s16 convolution function. |
| arm_convolve_s16_group_ch_mult_1Convolution | s16 grouped convolution optimized for the case where filterdims->c == 1 and inputch == outputch (channel multiplier = 1). |
| arm_convolve_s4Convolution | Basic s4 convolution function. |
| arm_convolve_s4_get_buffer_sizeConvolution | Get the required buffer size for s4 convolution function. |
| arm_convolve_s8Convolution | Basic s8 convolution function. |
| arm_convolve_s8_3x3_c16_s1Convolution | s8 3x3 convolution over 16 input channels with unit stride. |
| arm_convolve_s8_get_buffer_sizeConvolution | Get the required buffer size for s8 convolution function. |
| arm_convolve_s8_get_buffer_size_mveConvolution | Get the required buffer size for armconvolves8 for Arm(R) Helium Architecture case. |
| arm_convolve_s8_get_weights_sum_sizeConvolution | Get the required buffer size for s8 convolution and depthwise convolution weight sum. |
| arm_convolve_s8_small_cinConvolution | s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model. |
| arm_convolve_weight_sumConvolution | Pre-computes per-output-channel weight sums for a standard convolution. |
| arm_convolve_wrapper_f16Convolution | Convolution wrapper using the CMSIS-NN baseline path. |
| arm_convolve_wrapper_f16_acc16Convolution | Convolution wrapper using the CMSIS-NN baseline path. |
| arm_convolve_wrapper_f16_get_buffer_sizeConvolution | Get the buffer size required by the convolution wrapper. |
| arm_convolve_wrapper_f32Convolution | Convolution wrapper using the CMSIS-NN baseline path. |
| arm_convolve_wrapper_f32_get_buffer_sizeConvolution | Get the buffer size required by the convolution wrapper. |
| arm_convolve_wrapper_s16Convolution | s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution. |
| arm_convolve_wrapper_s16_get_buffer_sizeConvolution | Get the required buffer size for armconvolvewrappers16. |
| arm_convolve_wrapper_s16_get_buffer_size_dspConvolution | Get the required buffer size for armconvolvewrappers16 for for processors with DSP extension. |
| arm_convolve_wrapper_s16_get_buffer_size_mveConvolution | Get the required buffer size for armconvolvewrappers16 for Arm(R) Helium Architecture case. |
| arm_convolve_wrapper_s4Convolution | s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution. |
| arm_convolve_wrapper_s4_get_buffer_sizeConvolution | Get the required buffer size for armconvolvewrappers4. |
| arm_convolve_wrapper_s4_get_buffer_size_dspConvolution | Get the required buffer size for armconvolvewrappers4 for processors with DSP extension. |
| arm_convolve_wrapper_s4_get_buffer_size_mveConvolution | Get the required buffer size for armconvolvewrappers4 for Arm(R) Helium Architecture case. |
| arm_convolve_wrapper_s8Convolution | s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution. |
| arm_convolve_wrapper_s8_get_buffer_sizeConvolution | Get the required buffer size for armconvolvewrappers8. |
| arm_convolve_wrapper_s8_get_buffer_size_dspConvolution | Get the required buffer size for armconvolvewrappers8 for processors with DSP extension. |
| arm_convolve_wrapper_s8_get_buffer_size_mveConvolution | Get the required buffer size for armconvolvewrappers8 for Arm(R) Helium Architecture case. |
| arm_depth_to_space_s16Data movement | Depth to Space function for s16 data type. |
| arm_depth_to_space_s8Data movement | Depth to Space function for s8 data type. |
| arm_depthwise_conv_3x3_s8Convolution | Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below). |
| arm_depthwise_conv_f16Convolution | Depthwise convolution, dispatch by layout. |
| arm_depthwise_conv_f16_acc16Convolution | Depthwise convolution, dispatch by layout. |
| arm_depthwise_conv_f16_get_buffer_sizeConvolution | Get the temporary buffer size required by depthwise convolution. |
| arm_depthwise_conv_f32Convolution | Depthwise convolution, dispatch by layout. |
| arm_depthwise_conv_f32_get_buffer_sizeConvolution | Get the temporary buffer size required by depthwise convolution. |
| arm_depthwise_conv_fast_s16Convolution | Optimized s16 depthwise convolution function with constraint that inchannel equals outchannel. |
| arm_depthwise_conv_fast_s16_get_buffer_sizeConvolution | Get the required buffer size for optimized s16 depthwise convolution function with constraint that inchannel equals outchannel. |
| arm_depthwise_conv_s16Convolution | Basic s16 depthwise convolution function that doesn't have any constraints on the input dimensions. |
| arm_depthwise_conv_s4Convolution | Basic s4 depthwise convolution function that doesn't have any constraints on the input dimensions. |
| arm_depthwise_conv_s4_optConvolution | Optimized s4 depthwise convolution function with constraint that inchannel equals outchannel. |
| arm_depthwise_conv_s4_opt_get_buffer_sizeConvolution | Get the required buffer size for optimized s4 depthwise convolution function with constraint that inchannel equals outchannel. |
| arm_depthwise_conv_s8Convolution | Basic s8 depthwise convolution function that doesn't have any constraints on the input dimensions. |
| arm_depthwise_conv_s8_optConvolution | Optimized s8 depthwise convolution function with constraint that inchannel equals outchannel. |
| arm_depthwise_conv_s8_opt_3x3Convolution | s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of armdepthwiseconvs8opt(), for layers in its gate. |
| arm_depthwise_conv_s8_opt_3x3_c64_s1Convolution | armdepthwiseconvs8opt3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps. |
| arm_depthwise_conv_s8_opt_3x3_get_buffer_sizeConvolution | Get the scratch size in bytes of armdepthwiseconvs8opt3x3() and armdepthwiseconvs8opt3x3c64s1(). |
| arm_depthwise_conv_s8_opt_channelwiseConvolution | The channel-vectorized path of armdepthwiseconvs8opt() on its own, without the planar attempt. |
| arm_depthwise_conv_s8_opt_get_buffer_sizeConvolution | Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel. |
| arm_depthwise_conv_s8_opt_planarConvolution | The planar path of armdepthwiseconvs8opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane. |
| arm_depthwise_conv_s8_opt_planar_supportedConvolution | Whether armdepthwiseconvs8opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels. |
| arm_depthwise_conv_wrapper_f16Convolution | Depthwise convolution wrapper using the CMSIS-NN baseline path. |
| arm_depthwise_conv_wrapper_f16_acc16Convolution | Depthwise convolution wrapper using the CMSIS-NN baseline path. |
| arm_depthwise_conv_wrapper_f16_get_buffer_sizeConvolution | Get the buffer size required by the depthwise convolution wrapper. |
| arm_depthwise_conv_wrapper_f32Convolution | Depthwise convolution wrapper using the CMSIS-NN baseline path. |
| arm_depthwise_conv_wrapper_f32_get_buffer_sizeConvolution | Get the buffer size required by the depthwise convolution wrapper. |
| arm_depthwise_conv_wrapper_s16Convolution | Wrapper function to pick the right optimized s16 depthwise convolution function. |
| arm_depthwise_conv_wrapper_s16_get_buffer_sizeConvolution | Get size of additional buffer required by armdepthwiseconvwrappers16(). |
| arm_depthwise_conv_wrapper_s16_get_buffer_size_dspConvolution | Get size of additional buffer required by armdepthwiseconvwrappers16() for processors with DSP extension. |
| arm_depthwise_conv_wrapper_s16_get_buffer_size_mveConvolution | Get size of additional buffer required by armdepthwiseconvwrappers16() for Arm(R) Helium Architecture case. |
| arm_depthwise_conv_wrapper_s4Convolution | Wrapper function to pick the right optimized s4 depthwise convolution function. |
| arm_depthwise_conv_wrapper_s4_get_buffer_sizeConvolution | Get size of additional buffer required by armdepthwiseconvwrappers4(). |
| arm_depthwise_conv_wrapper_s4_get_buffer_size_dspConvolution | Get size of additional buffer required by armdepthwiseconvwrappers4() for processors with DSP extension. |
| arm_depthwise_conv_wrapper_s4_get_buffer_size_mveConvolution | Get size of additional buffer required by armdepthwiseconvwrappers4() for Arm(R) Helium Architecture case. |
| arm_depthwise_conv_wrapper_s8Convolution | Wrapper function to pick the right optimized s8 depthwise convolution function. |
| arm_depthwise_conv_wrapper_s8_get_buffer_sizeConvolution | Get size of additional buffer required by armdepthwiseconvwrappers8(). |
| arm_depthwise_conv_wrapper_s8_get_buffer_size_dspConvolution | Get size of additional buffer required by armdepthwiseconvwrappers8() for processors with DSP extension. |
| arm_depthwise_conv_wrapper_s8_get_buffer_size_mveConvolution | Get size of additional buffer required by armdepthwiseconvwrappers8() for Arm(R) Helium Architecture case. |
| arm_depthwise_convolve_weight_sumConvolution | Pre-computes per-channel weight sums for a depthwise convolution. |
| arm_depthwise_nhwc_conv_f16Convolution | Depthwise convolution, NHWC layout. |
| arm_depthwise_nhwc_conv_f16_acc16Convolution | Depthwise convolution, NHWC layout. |
| arm_depthwise_nhwc_conv_f32Convolution | Depthwise convolution, NHWC layout. |
| arm_dequantize_f16_f32Pooling, softmax, quantization | Widen a float16 vector to float32. |
| arm_dequantize_s16_f32Pooling, softmax, quantization | Dequantize an int16t array back to floating-point format. |
| arm_dequantize_s8_f32Pooling, softmax, quantization | Dequantize an int8t array back to floating-point format. |
| arm_dynamic_update_slice_s16Data movement | Update a slice of an int16 operand tensor at runtime-determined indices. |
| arm_dynamic_update_slice_s8Data movement | Update a slice of an int8 operand tensor at runtime-determined indices. |
| arm_elementwise_add_broadcast_f16Elementwise | Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp. |
| arm_elementwise_add_broadcast_f32Elementwise | Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp. |
| arm_elementwise_add_f16Elementwise | Elementwise add with optional output clamp. |
| arm_elementwise_add_f32Elementwise | Elementwise add with optional output clamp. |
| arm_elementwise_add_fp16Elementwise | Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate armelementwiseaddf16(). |
| arm_elementwise_add_s16Elementwise | s16 elementwise add of two vectors |
| arm_elementwise_add_s8Elementwise | s8 elementwise add of two vectors |
| arm_elementwise_mul_broadcast_f16Elementwise | Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp. |
| arm_elementwise_mul_broadcast_f32Elementwise | Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp. |
| arm_elementwise_mul_f16Elementwise | Elementwise multiply with optional output clamp. |
| arm_elementwise_mul_f32Elementwise | Elementwise multiply with optional output clamp. |
| arm_elementwise_mul_s16Elementwise | s16 elementwise multiplication |
| arm_elementwise_mul_s8Elementwise | s8 elementwise multiplication |
| arm_elementwise_prelu_s16Elementwise | Elementwise S16 PReLU activation function. |
| arm_elementwise_prelu_s8Elementwise | Elementwise S8 PReLU activation function. |
| arm_elementwise_squared_difference_f16Elementwise | Elementwise squared difference of two float16 vectors. |
| arm_elementwise_squared_difference_s16Elementwise | s16 elementwise squared difference of two vectors. |
| arm_elementwise_squared_difference_s8Elementwise | s8 elementwise squared difference of two vectors. |
| arm_elementwise_sub_broadcast_f16Elementwise | Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp. |
| arm_elementwise_sub_broadcast_f32Elementwise | Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp. |
| arm_elementwise_sub_f16Elementwise | Elementwise subtract with optional output clamp. |
| arm_elementwise_sub_f32Elementwise | Elementwise subtract with optional output clamp. |
| arm_elementwise_sub_s16Elementwise | s16 elementwise subtract of two vectors |
| arm_elementwise_sub_s8Elementwise | s8 elementwise subtract of two vectors |
| arm_equal_s16Reduction and comparison | s16 elementwise equality comparison with support for broadcasting. |
| arm_equal_s8Reduction and comparison | s8 elementwise equality comparison with support for broadcasting. |
| arm_fully_connected_f16Fully connected | Fully connected layer, dispatch by layout. |
| arm_fully_connected_f16_acc16Fully connected | Fully connected layer, dispatch by layout. |
| arm_fully_connected_f16_get_buffer_sizeFully connected | Get the temporary buffer size required by the fully connected layer. |
| arm_fully_connected_f32Fully connected | Fully connected layer, dispatch by layout. |
| arm_fully_connected_f32_get_buffer_sizeFully connected | Get the temporary buffer size required by the fully connected layer. |
| arm_fully_connected_nhwc_f16Fully connected | Fully connected layer, NHWC layout. |
| arm_fully_connected_nhwc_f16_acc16Fully connected | Fully connected layer, NHWC layout. |
| arm_fully_connected_nhwc_f32Fully connected | Fully connected layer, NHWC layout. |
| arm_fully_connected_per_channel_s16Fully connected | Basic s16 Fully Connected function using per channel quantization. |
| arm_fully_connected_per_channel_s16_get_buffer_sizeFully connected | Get size of additional buffer required by armfullyconnectedperchannels16(). |
| arm_fully_connected_per_channel_s16_get_buffer_size_dspFully connected | Get size of additional buffer required by armfullyconnectedperchannels16() for processors with DSP extension. |
| arm_fully_connected_per_channel_s16_get_buffer_size_mveFully connected | Get size of additional buffer required by armfullyconnectedperchannels16() for Arm(R) Helium Architecture case. |
| arm_fully_connected_per_channel_s8Fully connected | Basic s8 Fully Connected function using per channel quantization. |
| arm_fully_connected_s16Fully connected | Basic s16 Fully Connected function. |
| arm_fully_connected_s16_get_buffer_sizeFully connected | Get size of additional buffer required by armfullyconnecteds16(). |
| arm_fully_connected_s16_get_buffer_size_dspFully connected | Get size of additional buffer required by armfullyconnecteds16() for processors with DSP extension. |
| arm_fully_connected_s16_get_buffer_size_mveFully connected | Get size of additional buffer required by armfullyconnecteds16() for Arm(R) Helium Architecture case. |
| arm_fully_connected_s4Fully connected | Basic s4 Fully Connected function. |
| arm_fully_connected_s8Fully connected | Basic s8 Fully Connected function. |
| arm_fully_connected_s8_get_buffer_sizeFully connected | Get size of additional buffer required by armfullyconnecteds8(). |
| arm_fully_connected_s8_get_buffer_size_dspFully connected | Get size of additional buffer required by armfullyconnecteds8() for processors with DSP extension. |
| arm_fully_connected_s8_get_buffer_size_mveFully connected | Get size of additional buffer required by armfullyconnecteds8() for Arm(R) Helium Architecture case. |
| arm_fully_connected_wrapper_s16Fully connected | s16 Fully Connected layer wrapper function |
| arm_fully_connected_wrapper_s8Fully connected | s8 Fully Connected layer wrapper function |
| arm_gather_f16Data movement | Gather contiguous slices along an axis. |
| arm_gather_f32Data movement | Gather contiguous slices along an axis. |
| arm_gather_nd_f16Data movement | Gather contiguous slices using coordinate tuples. |
| arm_gather_nd_f32Data movement | Gather contiguous slices using coordinate tuples. |
| arm_gather_nd_s16Data movement | Gathernd slices for int16 tensors. |
| arm_gather_nd_s8Data movement | Gathernd slices for int8 tensors. |
| arm_gather_s16Data movement | Gather elements along an axis for int16 tensors. |
| arm_gather_s8Data movement | Gather elements along an axis for int8 tensors. |
| arm_greater_equal_s16Reduction and comparison | s16 elementwise greater-or-equal comparison with support for broadcasting. |
| arm_greater_equal_s8Reduction and comparison | s8 elementwise greater-or-equal comparison with support for broadcasting. |
| arm_greater_s16Reduction and comparison | s16 elementwise greater-than comparison with support for broadcasting. |
| arm_greater_s8Reduction and comparison | s8 elementwise greater-than comparison with support for broadcasting. |
| arm_gru_unidirectional_f16Sequence | Unidirectional GRU layer for float16 input, output and state. |
| arm_gru_unidirectional_f16_temp1_get_buffer_sizeSequence | Get size of the temp1 scratch buffer required by armgruunidirectionalf16(). |
| arm_gru_unidirectional_f32Sequence | Unidirectional GRU layer for float32 input, output and state. |
| arm_gru_unidirectional_f32_temp1_get_buffer_sizeSequence | Get size of the temp1 scratch buffer required by armgruunidirectionalf32(). |
| arm_hard_swish_compat_s8Activation | S8 Hard-Swish activation function (compatibility version). |
| arm_hard_swish_f16Activation | Hard swish activation for float16 data. |
| arm_hard_swish_f32Activation | Hard swish activation for float32 data. |
| arm_hard_swish_precise_s16Activation | S16 Hard-Swish activation function (precise version). |
| arm_hard_swish_precise_s8Activation | S8 Hard-Swish activation function (precise version). |
| arm_leaky_relu_s16Activation | S16 Leaky ReLU activation function. |
| arm_leaky_relu_s8Activation | S8 Leaky ReLU activation function. |
| arm_less_equal_s16Reduction and comparison | s16 elementwise less-or-equal comparison with support for broadcasting. |
| arm_less_equal_s8Reduction and comparison | s8 elementwise less-or-equal comparison with support for broadcasting. |
| arm_less_s16Reduction and comparison | s16 elementwise less-than comparison with support for broadcasting. |
| arm_less_s8Reduction and comparison | s8 elementwise less-than comparison with support for broadcasting. |
| arm_logistic_s16Activation | Logistic activation function for s16. |
| arm_lstm_unidirectional_f16Sequence | Unidirectional LSTM inference. |
| arm_lstm_unidirectional_f16_temp1_get_buffer_sizeSequence | Get size of the temp1 scratch buffer required by armlstmunidirectionalf16(). |
| arm_lstm_unidirectional_f16_temp2_get_buffer_sizeSequence | Get size of the temp2 scratch buffer required by armlstmunidirectionalf16(). |
| arm_lstm_unidirectional_f32Sequence | Unidirectional LSTM inference. |
| arm_lstm_unidirectional_f32_temp1_get_buffer_sizeSequence | Get size of the temp1 scratch buffer required by armlstmunidirectionalf32(). |
| arm_lstm_unidirectional_f32_temp2_get_buffer_sizeSequence | Get size of the temp2 scratch buffer required by armlstmunidirectionalf32(). |
| arm_lstm_unidirectional_s16Sequence | LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias. |
| arm_lstm_unidirectional_s16_temp1_get_buffer_sizeSequence | Get size of the temp1 scratch buffer required by armlstmunidirectionals16(). |
| arm_lstm_unidirectional_s16_temp2_get_buffer_sizeSequence | Get size of the temp2 scratch buffer required by armlstmunidirectionals16(). |
| arm_lstm_unidirectional_s8Sequence | LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias. |
| arm_lstm_unidirectional_s8_temp1_get_buffer_sizeSequence | Get size of the temp1 scratch buffer required by armlstmunidirectionals8(). |
| arm_lstm_unidirectional_s8_temp2_get_buffer_sizeSequence | Get size of the temp2 scratch buffer required by armlstmunidirectionals8(). |
| arm_max_pool_f16Pooling, softmax, quantization | Max pooling. |
| arm_max_pool_f32Pooling, softmax, quantization | Max pooling. |
| arm_max_pool_s16Pooling, softmax, quantization | s16 max pooling function. |
| arm_max_pool_s8Pooling, softmax, quantization | s8 max pooling function. |
| arm_maximum_f16Elementwise | Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. |
| arm_maximum_f32Elementwise | Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. |
| arm_maximum_s16Elementwise | s16 elementwise maximum w/ support for broadcasting and scalar inputs. |
| arm_maximum_s8Elementwise | s8 elementwise maximum w/ support for broadcasting and scalar inputs. |
| arm_mean_s16Reduction and comparison | Computes the mean of the input tensor along the specified axis. |
| arm_mean_s8Reduction and comparison | Computes the mean of the input tensor along the specified axis. |
| arm_minimum_f16Elementwise | Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. |
| arm_minimum_f32Elementwise | Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. |
| arm_minimum_s16Elementwise | s16 elementwise minimum w/ support for broadcasting and scalar inputs. |
| arm_minimum_s8Elementwise | s8 elementwise minimum w/ support for broadcasting and scalar inputs. |
| arm_mirror_pad_s16Data movement | Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode). |
| arm_mirror_pad_s8Data movement | Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode). |
| arm_mul_s16Elementwise | s16 elementwise multiplication of two tensors with support for broadcasting. |
| arm_mul_s8Elementwise | s8 elementwise multiplication of two tensors with support for broadcasting. |
| arm_mul_scalar_s16Elementwise | s16 elementwise multiplication of scalar and vector |
| arm_mul_scalar_s8Elementwise | s8 elementwise multiplication of scalar and vector |
| arm_nn_abs_f16Elementwise | Elementwise absolute value. |
| arm_nn_abs_f32Elementwise | Elementwise absolute value. |
| arm_nn_activation_f16Activation | Elementwise activation. |
| arm_nn_activation_f32Activation | Elementwise activation. |
| arm_nn_activation_s16Activation | s16 neural network activation function using direct table look-up |
| arm_nn_fill_f16Data movement | Fill a float16 vector with one value; bit copy of value, NaN payload included. |
| arm_nn_fill_f32Data movement | Fill a float32 vector with one value. |
| arm_nn_mean_f16Reduction and comparison | Computes the mean of a float16 tensor along the specified axes. |
| arm_nn_mean_f32Reduction and comparison | Computes the mean of a float32 tensor along the specified axes. |
| arm_nn_sqrt_f16Elementwise | Elementwise square root of a float16 tensor. |
| arm_nn_sqrt_f32Elementwise | Elementwise square root. |
| arm_not_equal_s16Reduction and comparison | s16 elementwise inequality comparison with support for broadcasting. |
| arm_not_equal_s8Reduction and comparison | s8 elementwise inequality comparison with support for broadcasting. |
| arm_pack_f16Data movement | Stack float32 tensors of equal shape along a new axis (TFLite PACK). |
| arm_pack_f32Data movement | Stack float32 tensors of equal shape along a new axis (TFLite PACK). |
| arm_pad_f16Data movement | Pad a tensor with a constant value. |
| arm_pad_f32Data movement | Pad a tensor with a constant value. |
| arm_pad_s16Data movement | Expands the size of the input by adding constant values before and after the data, in all dimensions. |
| arm_pad_s8Data movement | Expands the size of the input by adding constant values before and after the data, in all dimensions. |
| arm_prelu_f16Activation | Parametric ReLU for float32 data. |
| arm_prelu_f32Activation | Parametric ReLU for float32 data. |
| arm_prelu_s16Activation | S16 PReLU activation function. |
| arm_prelu_s8Activation | S8 PReLU activation function. |
| arm_prelu_scalar_s16Activation | Scalar S16 PReLU activation function. |
| arm_prelu_scalar_s8Activation | Scalar S8 PReLU activation function. |
| arm_quantize_f32_s16Pooling, softmax, quantization | Quantize a floating-point array into int16t format. |
| arm_quantize_f32_s8Pooling, softmax, quantization | Quantize a floating-point array into int8t format. |
| arm_reduce_max_f16Reduction and comparison | Reduces a f16 NHWC tensor to its maximum along a binary axis mask. |
| arm_reduce_max_f32Reduction and comparison | Reduces a f32 NHWC tensor to its maximum along a binary axis mask. |
| arm_reduce_max_s16Reduction and comparison | Computes the max of the input tensor along the specified axis. |
| arm_reduce_max_s8Reduction and comparison | Computes the max of the input tensor along the specified axis. |
| arm_reduce_min_f16Reduction and comparison | Reduces a f16 NHWC tensor to its minimum along a binary axis mask. |
| arm_reduce_min_f32Reduction and comparison | Reduces a f32 NHWC tensor to its minimum along a binary axis mask. |
| arm_reduce_min_s16Reduction and comparison | Computes the min of the input tensor along the specified axis. |
| arm_reduce_min_s8Reduction and comparison | Computes the min of the input tensor along the specified axis. |
| arm_reduce_sum_f16Reduction and comparison | Computes the sum of the input tensor along the specified axes. |
| arm_reduce_sum_f32Reduction and comparison | Computes the sum of the input tensor along the specified axes. |
| arm_relu_generic_s16Activation | S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc. |
| arm_relu_generic_s8Activation | S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc. |
| arm_relu_q15Activation | Q15 RELU function. |
| arm_relu_q7Activation | Q7 RELU function. |
| arm_relu_s16Activation | S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767. |
| arm_relu_s8Activation | S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127. |
| arm_relu6_q7Activation | Q7 RELU6 function. |
| arm_requantize_s16_s16Pooling, softmax, quantization | Requantize an int16t array to another int16t range with a different scale. |
| arm_requantize_s8_s8Pooling, softmax, quantization | Requantize an int8t array to another int8t range with a different scale. |
| arm_reshape_f16Data movement | Reshape by copying data without changing element order. |
| arm_reshape_f32Data movement | Reshape by copying data without changing element order. |
| arm_reshape_s8Data movement | Reshape a s8 vector into another with different shape. |
| arm_resize_nearest_neighbor_f16Data movement | Nearest-neighbor resize of a float32 NHWC tensor. |
| arm_resize_nearest_neighbor_f16_get_buffer_sizeData movement | Scratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16(). |
| arm_resize_nearest_neighbor_f32Data movement | Nearest-neighbor resize of a float32 NHWC tensor. |
| arm_resize_nearest_neighbor_f32_get_buffer_sizeData movement | Scratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16(). |
| arm_resize_nearest_neighbor_s16Data movement | Nearest neighbor resize function for s16 data. |
| arm_resize_nearest_neighbor_s8Data movement | Nearest neighbor resize function for s8 data. |
| arm_reverse_sequence_s16Data movement | Reverse variable-length sequences along a dimension for int16. |
| arm_reverse_sequence_s8Data movement | Reverse variable-length sequences along a dimension for int8. |
| arm_rsqrt_f16Elementwise | Elementwise reciprocal square root of a float16 tensor, 1 / sqrt(x). |
| arm_rsqrt_f32Elementwise | Elementwise reciprocal square root, 1 / sqrt(x). |
| arm_rsqrt_s16_per_opElementwise | INT16 reciprocal square root using a per-operator LUT. |
| arm_rsqrt_s16_universalElementwise | INT16 reciprocal square root using a shared universal LUT. |
| arm_scatter_nd_s16Data movement | Scatter updates into a zero-initialized output tensor for int16. |
| arm_scatter_nd_s8Data movement | Scatter updates into a zero-initialized output tensor for int8. |
| arm_select_v2_s16Elementwise | SELECTV2 with broadcast for int16 tensors. |
| arm_select_v2_s8Elementwise | SELECTV2 with broadcast for int8 tensors. |
| arm_softmax_f16Pooling, softmax, quantization | Softmax using the float-native API signature. |
| arm_softmax_f32Pooling, softmax, quantization | Softmax using the float-native API signature. |
| arm_softmax_s16Pooling, softmax, quantization | S16 softmax function. |
| arm_softmax_s8Pooling, softmax, quantization | S8 softmax function. |
| arm_softmax_s8_s16Pooling, softmax, quantization | S8 to s16 softmax function. |
| arm_softmax_u8Pooling, softmax, quantization | U8 softmax function. |
| arm_space_to_batch_nd_s16Data movement | Space to Batch ND function for s16 data type. |
| arm_space_to_batch_nd_s8Data movement | Space to Batch ND function for s8 data type. |
| arm_space_to_depth_s16Data movement | Space to Depth function for s16 data type. |
| arm_space_to_depth_s8Data movement | Space to Depth function for s8 data type. |
| arm_split_f16Data movement | Split a float32 tensor of any rank into several tensors along one axis. |
| arm_split_f32Data movement | Split a float32 tensor of any rank into several tensors along one axis. |
| arm_split_s16Data movement | int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis |
| arm_split_s8Data movement | int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis |
| arm_sqrt_s16Elementwise | s16 elementwise square root using piecewise LUT with linear interpolation |
| arm_sqrt_s16_tablefreeElementwise | s16 elementwise square root without a lookup table |
| arm_sqrt_s8Elementwise | s8 elementwise square root |
| arm_squared_difference_s16Elementwise | s16 elementwise squared difference of two tensors with support for broadcasting. |
| arm_squared_difference_s8Elementwise | s8 elementwise squared difference of two tensors with support for broadcasting. |
| arm_squared_difference_scalar_s16Elementwise | s16 elementwise squared difference of scalar and vector. |
| arm_squared_difference_scalar_s8Elementwise | s8 elementwise squared difference of scalar and vector. |
| arm_strided_slice_f16Data movement | Strided slice for float32 data (pure copy, TensorFlow Lite compatible). |
| arm_strided_slice_f32Data movement | Strided slice for float32 data (pure copy, TensorFlow Lite compatible). |
| arm_strided_slice_s16Data movement | Strided slice function for int16 data. |
| arm_strided_slice_s32Data movement | Strided slice function for int32 data. |
| arm_strided_slice_s8Data movement | Strided slice function for int8 data. |
| arm_sub_s16Elementwise | s16 elementwise subtraction of two tensors with support for broadcasting. |
| arm_sub_s8Elementwise | s8 elementwise subtraction of two tensors with support for broadcasting. |
| arm_sub_scalar_s16Elementwise | s16 elementwise subtract of scalar and vector (scalar - vector) |
| arm_sub_scalar_s8Elementwise | s8 elementwise subtract of scalar and vector (scalar - vector) |
| arm_svdf_f16Sequence | Stateful singular value decomposition filter, float16 variant. |
| arm_svdf_f16_input_ctx_get_buffer_sizeSequence | Get size of the inputctx staging buffer required by armsvdff16(). |
| arm_svdf_f16_output_ctx_get_buffer_sizeSequence | Get size of the outputctx staging buffer required by armsvdff16(). |
| arm_svdf_f32Sequence | Stateful singular value decomposition filter. |
| arm_svdf_f32_input_ctx_get_buffer_sizeSequence | Get size of the inputctx staging buffer required by armsvdff32(). |
| arm_svdf_f32_output_ctx_get_buffer_sizeSequence | Get size of the outputctx staging buffer required by armsvdff32(). |
| arm_svdf_s8Sequence | s8 SVDF function with 8 bit state tensor and 8 bit time weights |
| arm_svdf_s8_get_buffer_sizeSequence | Get size of the kernel-sum buffer required by armsvdfs8(). |
| arm_svdf_s8_get_buffer_size_dspSequence | Get size of the kernel-sum buffer required by armsvdfs8() for processors with DSP extension. |
| arm_svdf_s8_get_buffer_size_mveSequence | Get size of the kernel-sum buffer required by armsvdfs8() for Arm(R) Helium Architecture case. |
| arm_svdf_s8_input_ctx_get_buffer_sizeSequence | Get size of the inputctx staging buffer required by armsvdfs8(). |
| arm_svdf_s8_output_ctx_get_buffer_sizeSequence | Get size of the outputctx staging buffer required by armsvdfs8(). |
| arm_svdf_state_s16_s8Sequence | s8 SVDF function with 16 bit state tensor and 16 bit time weights |
| arm_svdf_state_s16_s8_input_ctx_get_buffer_sizeSequence | Get size of the inputctx staging buffer required by armsvdfstates16s8(). |
| arm_svdf_state_s16_s8_output_ctx_get_buffer_sizeSequence | Get size of the outputctx staging buffer required by armsvdfstates16s8(). |
| arm_tanh_s16Activation | Tanh activation function for s16. |
| arm_tile_s16Data movement | Tile an int16 tensor along each dimension. |
| arm_tile_s8Data movement | Tile an int8 tensor along each dimension. |
| arm_transpose_conv_f16Convolution | Transpose convolution, dispatch by layout. |
| arm_transpose_conv_f16_get_buffer_sizeConvolution | Get the temporary buffer size required by transpose convolution. |
| arm_transpose_conv_f16_get_reverse_conv_buffer_sizeConvolution | Get the reverse-convolution workspace size used by transpose convolution helpers. |
| arm_transpose_conv_f32Convolution | Transpose convolution, dispatch by layout. |
| arm_transpose_conv_f32_get_buffer_sizeConvolution | Get the temporary buffer size required by transpose convolution. |
| arm_transpose_conv_f32_get_reverse_conv_buffer_sizeConvolution | Get the reverse-convolution workspace size used by transpose convolution helpers. |
| arm_transpose_conv_nhwc_f16Convolution | Transpose convolution, NHWC layout. |
| arm_transpose_conv_nhwc_f32Convolution | Transpose convolution, NHWC layout. |
| arm_transpose_conv_s8Convolution | Basic s8 transpose convolution function. |
| arm_transpose_conv_s8_get_buffer_sizeConvolution | Get the required buffer size for ctx in s8 transpose conv function. |
| arm_transpose_conv_s8_get_buffer_size_mveConvolution | Get size of additional buffer required by armtransposeconvs8() for Arm(R) Helium Architecture case. |
| arm_transpose_conv_s8_get_reverse_conv_buffer_sizeConvolution | Get the required buffer size for outputctx in s8 transpose conv function. |
| arm_transpose_conv_wrapper_f16Convolution | Transpose convolution wrapper using the CMSIS-NN baseline path. |
| arm_transpose_conv_wrapper_f32Convolution | Transpose convolution wrapper using the CMSIS-NN baseline path. |
| arm_transpose_conv_wrapper_s8Convolution | Wrapper to select optimal transposed convolution algorithm depending on parameters. |
| arm_transpose_f16Data movement | Transpose a floating-point tensor. |
| arm_transpose_f32Data movement | Transpose a floating-point tensor. |
| arm_transpose_s16Data movement | Basic s16 transpose function. |
| arm_transpose_s8Data movement | Basic transpose function. |
| arm_unpack_f16Data movement | Unstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK). |
| arm_unpack_f32Data movement | Unstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK). |
| arm_vector_sum_s8Reduction and comparison | Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s32 biasdata. |
| arm_vector_sum_s8_s64Reduction and comparison | Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s64 biasdata. |
| arm_where_s16Reduction and comparison | WHERE operator: return coordinates of non-zero elements in condition (int16). |
| arm_where_s8Reduction and comparison | WHERE operator: return coordinates of non-zero elements in condition. |