Skip to content
heliaCORE
API reference
HELIA HUB

Kernel index

Search by function name or description, then narrow the results by operator group and data type. Select a function to read its parameters and requirements.

426 of 426 functions

Find a kernel
FunctionDescription
arm_abs_s16Elementwises16 elementwise absolute value
arm_abs_s8Elementwises8 elementwise absolute value
arm_add_s16Elementwises16 elementwise add of two tensors with support for broadcasting.
arm_add_s8Elementwises8 elementwise add of two tensors with support for broadcasting.
arm_add_scalar_s16Elementwises16 elementwise add of scalar and vector
arm_add_scalar_s8Elementwises8 elementwise add of scalar and vector
arm_argmax_f16Reduction and comparisonReturns the first maximum's axis-relative INT32 index for a f16 tensor.
arm_argmax_f32Reduction and comparisonReturns the first maximum's axis-relative INT32 index for a f32 tensor.
arm_argmax_s16Reduction and comparisonCompute ArgMax indices of an s16 tensor along a specific axis.
arm_argmax_s8Reduction and comparisonCompute ArgMax indices of an s8 tensor along a specific axis.
arm_argmin_f16Reduction and comparisonReturns the first minimum's axis-relative INT32 index for a f16 tensor.
arm_argmin_f32Reduction and comparisonReturns the first minimum's axis-relative INT32 index for a f32 tensor.
arm_argmin_s16Reduction and comparisonCompute ArgMin indices of an s16 tensor along a specific axis.
arm_argmin_s8Reduction and comparisonCompute ArgMin indices of an s8 tensor along a specific axis.
arm_avg_pool_f16Pooling, softmax, quantizationAverage pooling.
arm_avg_pool_f32Pooling, softmax, quantizationAverage pooling.
arm_avgpool_s16Pooling, softmax, quantizations16 average pooling function.
arm_avgpool_s16_get_buffer_sizePooling, softmax, quantizationGet the required buffer size for S16 average pooling function.
arm_avgpool_s16_get_buffer_size_dspPooling, softmax, quantizationGet the required buffer size for S16 average pooling function for processors with DSP extension.
arm_avgpool_s16_get_buffer_size_mvePooling, softmax, quantizationGet the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case.
arm_avgpool_s8Pooling, softmax, quantizations8 average pooling function.
arm_avgpool_s8_get_buffer_sizePooling, softmax, quantizationGet the required buffer size for S8 average pooling function.
arm_avgpool_s8_get_buffer_size_dspPooling, softmax, quantizationGet the required buffer size for S8 average pooling function for processors with DSP extension.
arm_avgpool_s8_get_buffer_size_mvePooling, softmax, quantizationGet the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case.
arm_batch_matmul_f16Fully connectedBatched matrix multiplication.
arm_batch_matmul_f16_get_buffer_sizeFully connectedGet the temporary buffer size required by batched matrix multiplication.
arm_batch_matmul_f32Fully connectedBatched matrix multiplication.
arm_batch_matmul_f32_get_buffer_sizeFully connectedGet the temporary buffer size required by batched matrix multiplication.
arm_batch_matmul_s16Fully connectedBatch matmul function with 16 bit input and output.
arm_batch_matmul_s8Fully connectedBatch matmul function with 8 bit input and output.
arm_batch_matmul_s8_get_buffer_sizeFully connectedGet size of the scratch buffer required by armbatchmatmuls8().
arm_batch_matmul_s8_get_buffer_size_dspFully connectedGet size of the scratch buffer required by armbatchmatmuls8() for processors with DSP extension.
arm_batch_matmul_s8_get_buffer_size_mveFully connectedGet size of the scratch buffer required by armbatchmatmuls8() for Arm(R) Helium Architecture case.
arm_batch_norm_f16ElementwiseApply batch normalization.
arm_batch_norm_f32ElementwiseApply batch normalization.
arm_batch_to_space_nd_s16Data movementBatch to Space ND function for s16 data type.
arm_batch_to_space_nd_s8Data movementBatch to Space ND function for s8 data type.
arm_broadcast_to_s16Data movementBroadcast an int16 tensor to a target shape.
arm_broadcast_to_s8Data movementBroadcast an int8 tensor to a target shape.
arm_clamp_s16ActivationS16 clamp function.
arm_clamp_s8ActivationS8 clamp function.
arm_comparison_s16Reduction and comparisons16 elementwise comparison with support for broadcasting.
arm_comparison_s8Reduction and comparisons8 elementwise comparison with support for broadcasting.
arm_concatenation_f16Data movementConcatenate float32 tensors of any rank along one axis.
arm_concatenation_f16_wData movementConcatenate tensors along the W axis.
arm_concatenation_f16_xData movementConcatenate tensors along the X axis.
arm_concatenation_f16_yData movementConcatenate tensors along the Y axis.
arm_concatenation_f16_zData movementConcatenate tensors along the Z axis.
arm_concatenation_f32Data movementConcatenate float32 tensors of any rank along one axis.
arm_concatenation_f32_wData movementConcatenate tensors along the W axis.
arm_concatenation_f32_xData movementConcatenate tensors along the X axis.
arm_concatenation_f32_yData movementConcatenate tensors along the Y axis.
arm_concatenation_f32_zData movementConcatenate tensors along the Z axis.
arm_concatenation_s16Data movementint16/uint16 concatenation function to be used for concatenating N-tensors along the target axis
arm_concatenation_s32Data movementint32/uint32 concatenation function to be used for concatenating N-tensors along the target axis
arm_concatenation_s8Data movementint8/uint8 concatenation function to be used for concatenating N-tensors along the target axis
arm_concatenation_s8_wData movementint8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to…
arm_concatenation_s8_xData movementint8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate.
arm_concatenation_s8_yData movementint8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate.
arm_concatenation_s8_zData movementint8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate.
arm_convolve_1_x_n_f16Convolution1xN convolution, dispatch by layout.
arm_convolve_1_x_n_f16_acc16Convolution1xN convolution, dispatch by layout.
arm_convolve_1_x_n_f16_get_buffer_sizeConvolutionGet the buffer size required by 1xN convolution.
arm_convolve_1_x_n_f32Convolution1xN convolution, dispatch by layout.
arm_convolve_1_x_n_f32_get_buffer_sizeConvolutionGet the buffer size required by 1xN convolution.
arm_convolve_1_x_n_nhwc_f16Convolution1xN convolution, NHWC layout.
arm_convolve_1_x_n_nhwc_f16_acc16Convolution1xN convolution, NHWC layout.
arm_convolve_1_x_n_nhwc_f32Convolution1xN convolution, NHWC layout.
arm_convolve_1_x_n_s4Convolution1xn convolution for s4 weights
arm_convolve_1_x_n_s4_get_buffer_sizeConvolutionGet the required additional buffer size for 1xn convolution.
arm_convolve_1_x_n_s8Convolution1xn convolution
arm_convolve_1_x_n_s8_get_buffer_sizeConvolutionGet the required additional buffer size for 1xn convolution.
arm_convolve_1x1_f16Convolution1x1 convolution, dispatch by layout.
arm_convolve_1x1_f16_acc16Convolution1x1 convolution, dispatch by layout.
arm_convolve_1x1_f16_get_buffer_sizeConvolutionGet the buffer size required by 1x1 convolution.
arm_convolve_1x1_f32Convolution1x1 convolution, dispatch by layout.
arm_convolve_1x1_f32_get_buffer_sizeConvolutionGet the buffer size required by 1x1 convolution.
arm_convolve_1x1_nhwc_f16Convolution1x1 convolution, NHWC layout.
arm_convolve_1x1_nhwc_f16_acc16Convolution1x1 convolution, NHWC layout.
arm_convolve_1x1_nhwc_f32Convolution1x1 convolution, NHWC layout.
arm_convolve_1x1_out_s8ConvolutionOptimised convolution for 1x1 output images (shape of BX1x1xCOUT) for 8x8 computations.
arm_convolve_1x1_out_s8_get_buffer_sizeConvolutionGet the required scratch buffer size for armconvolve1x1outs8().
arm_convolve_1x1_s16_ns_np_ndConvolutionPointwise s16 convolution function: no stride, no padding, no dilation.
arm_convolve_1x1_s4Convolutions4 version for 1x1 convolution with support for non-unity stride values
arm_convolve_1x1_s4_fastConvolutionFast s4 version for 1x1 convolution (non-square shape).
arm_convolve_1x1_s4_fast_get_buffer_sizeConvolutionGet the required buffer size for armconvolve1x1s4fast.
arm_convolve_1x1_s8Convolutions8 version for 1x1 convolution with support for non-unity stride values
arm_convolve_1x1_s8_fastConvolutionFast s8 version for 1x1 convolution (non-square shape).
arm_convolve_1x1_s8_fast_get_buffer_sizeConvolutionGet the required buffer size for armconvolve1x1s8fast.
arm_convolve_even_s4ConvolutionBasic s4 convolution function with a requirement of even number of kernels.
arm_convolve_even_s4_get_buffer_sizeConvolutionGet the required buffer size for armconvolveevens4.
arm_convolve_f16ConvolutionConvolution, dispatch by layout.
arm_convolve_f16_acc16ConvolutionConvolution, dispatch by layout.
arm_convolve_f16_get_buffer_sizeConvolutionGet the temporary buffer size required by convolution.
arm_convolve_f32ConvolutionConvolution, dispatch by layout.
arm_convolve_f32_get_buffer_sizeConvolutionGet the temporary buffer size required by convolution.
arm_convolve_nhwc_f16ConvolutionConvolution, NHWC layout.
arm_convolve_nhwc_f16_acc16ConvolutionConvolution, NHWC layout.
arm_convolve_nhwc_f32ConvolutionConvolution, NHWC layout.
arm_convolve_s16ConvolutionBasic s16 convolution function.
arm_convolve_s16_fast_small_kernelConvolutionarmconvolves16fastsmallkernel function.
arm_convolve_s16_get_buffer_sizeConvolutionGet the required buffer size for s16 convolution function.
arm_convolve_s16_group_ch_mult_1Convolutions16 grouped convolution optimized for the case where filterdims->c == 1 and inputch == outputch (channel multiplier = 1).
arm_convolve_s4ConvolutionBasic s4 convolution function.
arm_convolve_s4_get_buffer_sizeConvolutionGet the required buffer size for s4 convolution function.
arm_convolve_s8ConvolutionBasic s8 convolution function.
arm_convolve_s8_3x3_c16_s1Convolutions8 3x3 convolution over 16 input channels with unit stride.
arm_convolve_s8_get_buffer_sizeConvolutionGet the required buffer size for s8 convolution function.
arm_convolve_s8_get_buffer_size_mveConvolutionGet the required buffer size for armconvolves8 for Arm(R) Helium Architecture case.
arm_convolve_s8_get_weights_sum_sizeConvolutionGet the required buffer size for s8 convolution and depthwise convolution weight sum.
arm_convolve_s8_small_cinConvolutions8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model.
arm_convolve_weight_sumConvolutionPre-computes per-output-channel weight sums for a standard convolution.
arm_convolve_wrapper_f16ConvolutionConvolution wrapper using the CMSIS-NN baseline path.
arm_convolve_wrapper_f16_acc16ConvolutionConvolution wrapper using the CMSIS-NN baseline path.
arm_convolve_wrapper_f16_get_buffer_sizeConvolutionGet the buffer size required by the convolution wrapper.
arm_convolve_wrapper_f32ConvolutionConvolution wrapper using the CMSIS-NN baseline path.
arm_convolve_wrapper_f32_get_buffer_sizeConvolutionGet the buffer size required by the convolution wrapper.
arm_convolve_wrapper_s16Convolutions16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
arm_convolve_wrapper_s16_get_buffer_sizeConvolutionGet the required buffer size for armconvolvewrappers16.
arm_convolve_wrapper_s16_get_buffer_size_dspConvolutionGet the required buffer size for armconvolvewrappers16 for for processors with DSP extension.
arm_convolve_wrapper_s16_get_buffer_size_mveConvolutionGet the required buffer size for armconvolvewrappers16 for Arm(R) Helium Architecture case.
arm_convolve_wrapper_s4Convolutions4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
arm_convolve_wrapper_s4_get_buffer_sizeConvolutionGet the required buffer size for armconvolvewrappers4.
arm_convolve_wrapper_s4_get_buffer_size_dspConvolutionGet the required buffer size for armconvolvewrappers4 for processors with DSP extension.
arm_convolve_wrapper_s4_get_buffer_size_mveConvolutionGet the required buffer size for armconvolvewrappers4 for Arm(R) Helium Architecture case.
arm_convolve_wrapper_s8Convolutions8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.
arm_convolve_wrapper_s8_get_buffer_sizeConvolutionGet the required buffer size for armconvolvewrappers8.
arm_convolve_wrapper_s8_get_buffer_size_dspConvolutionGet the required buffer size for armconvolvewrappers8 for processors with DSP extension.
arm_convolve_wrapper_s8_get_buffer_size_mveConvolutionGet the required buffer size for armconvolvewrappers8 for Arm(R) Helium Architecture case.
arm_depth_to_space_s16Data movementDepth to Space function for s16 data type.
arm_depth_to_space_s8Data movementDepth to Space function for s8 data type.
arm_depthwise_conv_3x3_s8ConvolutionOptimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below).
arm_depthwise_conv_f16ConvolutionDepthwise convolution, dispatch by layout.
arm_depthwise_conv_f16_acc16ConvolutionDepthwise convolution, dispatch by layout.
arm_depthwise_conv_f16_get_buffer_sizeConvolutionGet the temporary buffer size required by depthwise convolution.
arm_depthwise_conv_f32ConvolutionDepthwise convolution, dispatch by layout.
arm_depthwise_conv_f32_get_buffer_sizeConvolutionGet the temporary buffer size required by depthwise convolution.
arm_depthwise_conv_fast_s16ConvolutionOptimized s16 depthwise convolution function with constraint that inchannel equals outchannel.
arm_depthwise_conv_fast_s16_get_buffer_sizeConvolutionGet the required buffer size for optimized s16 depthwise convolution function with constraint that inchannel equals outchannel.
arm_depthwise_conv_s16ConvolutionBasic s16 depthwise convolution function that doesn't have any constraints on the input dimensions.
arm_depthwise_conv_s4ConvolutionBasic s4 depthwise convolution function that doesn't have any constraints on the input dimensions.
arm_depthwise_conv_s4_optConvolutionOptimized s4 depthwise convolution function with constraint that inchannel equals outchannel.
arm_depthwise_conv_s4_opt_get_buffer_sizeConvolutionGet the required buffer size for optimized s4 depthwise convolution function with constraint that inchannel equals outchannel.
arm_depthwise_conv_s8ConvolutionBasic s8 depthwise convolution function that doesn't have any constraints on the input dimensions.
arm_depthwise_conv_s8_optConvolutionOptimized s8 depthwise convolution function with constraint that inchannel equals outchannel.
arm_depthwise_conv_s8_opt_3x3Convolutions8 3x3 depthwise convolution that reads the input in place, without the im2col copy of armdepthwiseconvs8opt(), for layers in its gate.
arm_depthwise_conv_s8_opt_3x3_c64_s1Convolutionarmdepthwiseconvs8opt3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps.
arm_depthwise_conv_s8_opt_3x3_get_buffer_sizeConvolutionGet the scratch size in bytes of armdepthwiseconvs8opt3x3() and armdepthwiseconvs8opt3x3c64s1().
arm_depthwise_conv_s8_opt_channelwiseConvolutionThe channel-vectorized path of armdepthwiseconvs8opt() on its own, without the planar attempt.
arm_depthwise_conv_s8_opt_get_buffer_sizeConvolutionGet the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel.
arm_depthwise_conv_s8_opt_planarConvolutionThe planar path of armdepthwiseconvs8opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane.
arm_depthwise_conv_s8_opt_planar_supportedConvolutionWhether armdepthwiseconvs8opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels.
arm_depthwise_conv_wrapper_f16ConvolutionDepthwise convolution wrapper using the CMSIS-NN baseline path.
arm_depthwise_conv_wrapper_f16_acc16ConvolutionDepthwise convolution wrapper using the CMSIS-NN baseline path.
arm_depthwise_conv_wrapper_f16_get_buffer_sizeConvolutionGet the buffer size required by the depthwise convolution wrapper.
arm_depthwise_conv_wrapper_f32ConvolutionDepthwise convolution wrapper using the CMSIS-NN baseline path.
arm_depthwise_conv_wrapper_f32_get_buffer_sizeConvolutionGet the buffer size required by the depthwise convolution wrapper.
arm_depthwise_conv_wrapper_s16ConvolutionWrapper function to pick the right optimized s16 depthwise convolution function.
arm_depthwise_conv_wrapper_s16_get_buffer_sizeConvolutionGet size of additional buffer required by armdepthwiseconvwrappers16().
arm_depthwise_conv_wrapper_s16_get_buffer_size_dspConvolutionGet size of additional buffer required by armdepthwiseconvwrappers16() for processors with DSP extension.
arm_depthwise_conv_wrapper_s16_get_buffer_size_mveConvolutionGet size of additional buffer required by armdepthwiseconvwrappers16() for Arm(R) Helium Architecture case.
arm_depthwise_conv_wrapper_s4ConvolutionWrapper function to pick the right optimized s4 depthwise convolution function.
arm_depthwise_conv_wrapper_s4_get_buffer_sizeConvolutionGet size of additional buffer required by armdepthwiseconvwrappers4().
arm_depthwise_conv_wrapper_s4_get_buffer_size_dspConvolutionGet size of additional buffer required by armdepthwiseconvwrappers4() for processors with DSP extension.
arm_depthwise_conv_wrapper_s4_get_buffer_size_mveConvolutionGet size of additional buffer required by armdepthwiseconvwrappers4() for Arm(R) Helium Architecture case.
arm_depthwise_conv_wrapper_s8ConvolutionWrapper function to pick the right optimized s8 depthwise convolution function.
arm_depthwise_conv_wrapper_s8_get_buffer_sizeConvolutionGet size of additional buffer required by armdepthwiseconvwrappers8().
arm_depthwise_conv_wrapper_s8_get_buffer_size_dspConvolutionGet size of additional buffer required by armdepthwiseconvwrappers8() for processors with DSP extension.
arm_depthwise_conv_wrapper_s8_get_buffer_size_mveConvolutionGet size of additional buffer required by armdepthwiseconvwrappers8() for Arm(R) Helium Architecture case.
arm_depthwise_convolve_weight_sumConvolutionPre-computes per-channel weight sums for a depthwise convolution.
arm_depthwise_nhwc_conv_f16ConvolutionDepthwise convolution, NHWC layout.
arm_depthwise_nhwc_conv_f16_acc16ConvolutionDepthwise convolution, NHWC layout.
arm_depthwise_nhwc_conv_f32ConvolutionDepthwise convolution, NHWC layout.
arm_dequantize_f16_f32Pooling, softmax, quantizationWiden a float16 vector to float32.
arm_dequantize_s16_f32Pooling, softmax, quantizationDequantize an int16t array back to floating-point format.
arm_dequantize_s8_f32Pooling, softmax, quantizationDequantize an int8t array back to floating-point format.
arm_dynamic_update_slice_s16Data movementUpdate a slice of an int16 operand tensor at runtime-determined indices.
arm_dynamic_update_slice_s8Data movementUpdate a slice of an int8 operand tensor at runtime-determined indices.
arm_elementwise_add_broadcast_f16ElementwiseElementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.
arm_elementwise_add_broadcast_f32ElementwiseElementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.
arm_elementwise_add_f16ElementwiseElementwise add with optional output clamp.
arm_elementwise_add_f32ElementwiseElementwise add with optional output clamp.
arm_elementwise_add_fp16ElementwiseLegacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate armelementwiseaddf16().
arm_elementwise_add_s16Elementwises16 elementwise add of two vectors
arm_elementwise_add_s8Elementwises8 elementwise add of two vectors
arm_elementwise_mul_broadcast_f16ElementwiseElementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.
arm_elementwise_mul_broadcast_f32ElementwiseElementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.
arm_elementwise_mul_f16ElementwiseElementwise multiply with optional output clamp.
arm_elementwise_mul_f32ElementwiseElementwise multiply with optional output clamp.
arm_elementwise_mul_s16Elementwises16 elementwise multiplication
arm_elementwise_mul_s8Elementwises8 elementwise multiplication
arm_elementwise_prelu_s16ElementwiseElementwise S16 PReLU activation function.
arm_elementwise_prelu_s8ElementwiseElementwise S8 PReLU activation function.
arm_elementwise_squared_difference_f16ElementwiseElementwise squared difference of two float16 vectors.
arm_elementwise_squared_difference_s16Elementwises16 elementwise squared difference of two vectors.
arm_elementwise_squared_difference_s8Elementwises8 elementwise squared difference of two vectors.
arm_elementwise_sub_broadcast_f16ElementwiseElementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.
arm_elementwise_sub_broadcast_f32ElementwiseElementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.
arm_elementwise_sub_f16ElementwiseElementwise subtract with optional output clamp.
arm_elementwise_sub_f32ElementwiseElementwise subtract with optional output clamp.
arm_elementwise_sub_s16Elementwises16 elementwise subtract of two vectors
arm_elementwise_sub_s8Elementwises8 elementwise subtract of two vectors
arm_equal_s16Reduction and comparisons16 elementwise equality comparison with support for broadcasting.
arm_equal_s8Reduction and comparisons8 elementwise equality comparison with support for broadcasting.
arm_fully_connected_f16Fully connectedFully connected layer, dispatch by layout.
arm_fully_connected_f16_acc16Fully connectedFully connected layer, dispatch by layout.
arm_fully_connected_f16_get_buffer_sizeFully connectedGet the temporary buffer size required by the fully connected layer.
arm_fully_connected_f32Fully connectedFully connected layer, dispatch by layout.
arm_fully_connected_f32_get_buffer_sizeFully connectedGet the temporary buffer size required by the fully connected layer.
arm_fully_connected_nhwc_f16Fully connectedFully connected layer, NHWC layout.
arm_fully_connected_nhwc_f16_acc16Fully connectedFully connected layer, NHWC layout.
arm_fully_connected_nhwc_f32Fully connectedFully connected layer, NHWC layout.
arm_fully_connected_per_channel_s16Fully connectedBasic s16 Fully Connected function using per channel quantization.
arm_fully_connected_per_channel_s16_get_buffer_sizeFully connectedGet size of additional buffer required by armfullyconnectedperchannels16().
arm_fully_connected_per_channel_s16_get_buffer_size_dspFully connectedGet size of additional buffer required by armfullyconnectedperchannels16() for processors with DSP extension.
arm_fully_connected_per_channel_s16_get_buffer_size_mveFully connectedGet size of additional buffer required by armfullyconnectedperchannels16() for Arm(R) Helium Architecture case.
arm_fully_connected_per_channel_s8Fully connectedBasic s8 Fully Connected function using per channel quantization.
arm_fully_connected_s16Fully connectedBasic s16 Fully Connected function.
arm_fully_connected_s16_get_buffer_sizeFully connectedGet size of additional buffer required by armfullyconnecteds16().
arm_fully_connected_s16_get_buffer_size_dspFully connectedGet size of additional buffer required by armfullyconnecteds16() for processors with DSP extension.
arm_fully_connected_s16_get_buffer_size_mveFully connectedGet size of additional buffer required by armfullyconnecteds16() for Arm(R) Helium Architecture case.
arm_fully_connected_s4Fully connectedBasic s4 Fully Connected function.
arm_fully_connected_s8Fully connectedBasic s8 Fully Connected function.
arm_fully_connected_s8_get_buffer_sizeFully connectedGet size of additional buffer required by armfullyconnecteds8().
arm_fully_connected_s8_get_buffer_size_dspFully connectedGet size of additional buffer required by armfullyconnecteds8() for processors with DSP extension.
arm_fully_connected_s8_get_buffer_size_mveFully connectedGet size of additional buffer required by armfullyconnecteds8() for Arm(R) Helium Architecture case.
arm_fully_connected_wrapper_s16Fully connecteds16 Fully Connected layer wrapper function
arm_fully_connected_wrapper_s8Fully connecteds8 Fully Connected layer wrapper function
arm_gather_f16Data movementGather contiguous slices along an axis.
arm_gather_f32Data movementGather contiguous slices along an axis.
arm_gather_nd_f16Data movementGather contiguous slices using coordinate tuples.
arm_gather_nd_f32Data movementGather contiguous slices using coordinate tuples.
arm_gather_nd_s16Data movementGathernd slices for int16 tensors.
arm_gather_nd_s8Data movementGathernd slices for int8 tensors.
arm_gather_s16Data movementGather elements along an axis for int16 tensors.
arm_gather_s8Data movementGather elements along an axis for int8 tensors.
arm_greater_equal_s16Reduction and comparisons16 elementwise greater-or-equal comparison with support for broadcasting.
arm_greater_equal_s8Reduction and comparisons8 elementwise greater-or-equal comparison with support for broadcasting.
arm_greater_s16Reduction and comparisons16 elementwise greater-than comparison with support for broadcasting.
arm_greater_s8Reduction and comparisons8 elementwise greater-than comparison with support for broadcasting.
arm_gru_unidirectional_f16SequenceUnidirectional GRU layer for float16 input, output and state.
arm_gru_unidirectional_f16_temp1_get_buffer_sizeSequenceGet size of the temp1 scratch buffer required by armgruunidirectionalf16().
arm_gru_unidirectional_f32SequenceUnidirectional GRU layer for float32 input, output and state.
arm_gru_unidirectional_f32_temp1_get_buffer_sizeSequenceGet size of the temp1 scratch buffer required by armgruunidirectionalf32().
arm_hard_swish_compat_s8ActivationS8 Hard-Swish activation function (compatibility version).
arm_hard_swish_f16ActivationHard swish activation for float16 data.
arm_hard_swish_f32ActivationHard swish activation for float32 data.
arm_hard_swish_precise_s16ActivationS16 Hard-Swish activation function (precise version).
arm_hard_swish_precise_s8ActivationS8 Hard-Swish activation function (precise version).
arm_leaky_relu_s16ActivationS16 Leaky ReLU activation function.
arm_leaky_relu_s8ActivationS8 Leaky ReLU activation function.
arm_less_equal_s16Reduction and comparisons16 elementwise less-or-equal comparison with support for broadcasting.
arm_less_equal_s8Reduction and comparisons8 elementwise less-or-equal comparison with support for broadcasting.
arm_less_s16Reduction and comparisons16 elementwise less-than comparison with support for broadcasting.
arm_less_s8Reduction and comparisons8 elementwise less-than comparison with support for broadcasting.
arm_logistic_s16ActivationLogistic activation function for s16.
arm_lstm_unidirectional_f16SequenceUnidirectional LSTM inference.
arm_lstm_unidirectional_f16_temp1_get_buffer_sizeSequenceGet size of the temp1 scratch buffer required by armlstmunidirectionalf16().
arm_lstm_unidirectional_f16_temp2_get_buffer_sizeSequenceGet size of the temp2 scratch buffer required by armlstmunidirectionalf16().
arm_lstm_unidirectional_f32SequenceUnidirectional LSTM inference.
arm_lstm_unidirectional_f32_temp1_get_buffer_sizeSequenceGet size of the temp1 scratch buffer required by armlstmunidirectionalf32().
arm_lstm_unidirectional_f32_temp2_get_buffer_sizeSequenceGet size of the temp2 scratch buffer required by armlstmunidirectionalf32().
arm_lstm_unidirectional_s16SequenceLSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias.
arm_lstm_unidirectional_s16_temp1_get_buffer_sizeSequenceGet size of the temp1 scratch buffer required by armlstmunidirectionals16().
arm_lstm_unidirectional_s16_temp2_get_buffer_sizeSequenceGet size of the temp2 scratch buffer required by armlstmunidirectionals16().
arm_lstm_unidirectional_s8SequenceLSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias.
arm_lstm_unidirectional_s8_temp1_get_buffer_sizeSequenceGet size of the temp1 scratch buffer required by armlstmunidirectionals8().
arm_lstm_unidirectional_s8_temp2_get_buffer_sizeSequenceGet size of the temp2 scratch buffer required by armlstmunidirectionals8().
arm_max_pool_f16Pooling, softmax, quantizationMax pooling.
arm_max_pool_f32Pooling, softmax, quantizationMax pooling.
arm_max_pool_s16Pooling, softmax, quantizations16 max pooling function.
arm_max_pool_s8Pooling, softmax, quantizations8 max pooling function.
arm_maximum_f16ElementwiseElementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape.
arm_maximum_f32ElementwiseElementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape.
arm_maximum_s16Elementwises16 elementwise maximum w/ support for broadcasting and scalar inputs.
arm_maximum_s8Elementwises8 elementwise maximum w/ support for broadcasting and scalar inputs.
arm_mean_s16Reduction and comparisonComputes the mean of the input tensor along the specified axis.
arm_mean_s8Reduction and comparisonComputes the mean of the input tensor along the specified axis.
arm_minimum_f16ElementwiseElementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape.
arm_minimum_f32ElementwiseElementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape.
arm_minimum_s16Elementwises16 elementwise minimum w/ support for broadcasting and scalar inputs.
arm_minimum_s8Elementwises8 elementwise minimum w/ support for broadcasting and scalar inputs.
arm_mirror_pad_s16Data movementMirror-pad an int16 tensor (REFLECT or SYMMETRIC mode).
arm_mirror_pad_s8Data movementMirror-pad an int8 tensor (REFLECT or SYMMETRIC mode).
arm_mul_s16Elementwises16 elementwise multiplication of two tensors with support for broadcasting.
arm_mul_s8Elementwises8 elementwise multiplication of two tensors with support for broadcasting.
arm_mul_scalar_s16Elementwises16 elementwise multiplication of scalar and vector
arm_mul_scalar_s8Elementwises8 elementwise multiplication of scalar and vector
arm_nn_abs_f16ElementwiseElementwise absolute value.
arm_nn_abs_f32ElementwiseElementwise absolute value.
arm_nn_activation_f16ActivationElementwise activation.
arm_nn_activation_f32ActivationElementwise activation.
arm_nn_activation_s16Activations16 neural network activation function using direct table look-up
arm_nn_fill_f16Data movementFill a float16 vector with one value; bit copy of value, NaN payload included.
arm_nn_fill_f32Data movementFill a float32 vector with one value.
arm_nn_mean_f16Reduction and comparisonComputes the mean of a float16 tensor along the specified axes.
arm_nn_mean_f32Reduction and comparisonComputes the mean of a float32 tensor along the specified axes.
arm_nn_sqrt_f16ElementwiseElementwise square root of a float16 tensor.
arm_nn_sqrt_f32ElementwiseElementwise square root.
arm_not_equal_s16Reduction and comparisons16 elementwise inequality comparison with support for broadcasting.
arm_not_equal_s8Reduction and comparisons8 elementwise inequality comparison with support for broadcasting.
arm_pack_f16Data movementStack float32 tensors of equal shape along a new axis (TFLite PACK).
arm_pack_f32Data movementStack float32 tensors of equal shape along a new axis (TFLite PACK).
arm_pad_f16Data movementPad a tensor with a constant value.
arm_pad_f32Data movementPad a tensor with a constant value.
arm_pad_s16Data movementExpands the size of the input by adding constant values before and after the data, in all dimensions.
arm_pad_s8Data movementExpands the size of the input by adding constant values before and after the data, in all dimensions.
arm_prelu_f16ActivationParametric ReLU for float32 data.
arm_prelu_f32ActivationParametric ReLU for float32 data.
arm_prelu_s16ActivationS16 PReLU activation function.
arm_prelu_s8ActivationS8 PReLU activation function.
arm_prelu_scalar_s16ActivationScalar S16 PReLU activation function.
arm_prelu_scalar_s8ActivationScalar S8 PReLU activation function.
arm_quantize_f32_s16Pooling, softmax, quantizationQuantize a floating-point array into int16t format.
arm_quantize_f32_s8Pooling, softmax, quantizationQuantize a floating-point array into int8t format.
arm_reduce_max_f16Reduction and comparisonReduces a f16 NHWC tensor to its maximum along a binary axis mask.
arm_reduce_max_f32Reduction and comparisonReduces a f32 NHWC tensor to its maximum along a binary axis mask.
arm_reduce_max_s16Reduction and comparisonComputes the max of the input tensor along the specified axis.
arm_reduce_max_s8Reduction and comparisonComputes the max of the input tensor along the specified axis.
arm_reduce_min_f16Reduction and comparisonReduces a f16 NHWC tensor to its minimum along a binary axis mask.
arm_reduce_min_f32Reduction and comparisonReduces a f32 NHWC tensor to its minimum along a binary axis mask.
arm_reduce_min_s16Reduction and comparisonComputes the min of the input tensor along the specified axis.
arm_reduce_min_s8Reduction and comparisonComputes the min of the input tensor along the specified axis.
arm_reduce_sum_f16Reduction and comparisonComputes the sum of the input tensor along the specified axes.
arm_reduce_sum_f32Reduction and comparisonComputes the sum of the input tensor along the specified axes.
arm_relu_generic_s16ActivationS16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.
arm_relu_generic_s8ActivationS8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.
arm_relu_q15ActivationQ15 RELU function.
arm_relu_q7ActivationQ7 RELU function.
arm_relu_s16ActivationS16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767.
arm_relu_s8ActivationS8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127.
arm_relu6_q7ActivationQ7 RELU6 function.
arm_requantize_s16_s16Pooling, softmax, quantizationRequantize an int16t array to another int16t range with a different scale.
arm_requantize_s8_s8Pooling, softmax, quantizationRequantize an int8t array to another int8t range with a different scale.
arm_reshape_f16Data movementReshape by copying data without changing element order.
arm_reshape_f32Data movementReshape by copying data without changing element order.
arm_reshape_s8Data movementReshape a s8 vector into another with different shape.
arm_resize_nearest_neighbor_f16Data movementNearest-neighbor resize of a float32 NHWC tensor.
arm_resize_nearest_neighbor_f16_get_buffer_sizeData movementScratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16().
arm_resize_nearest_neighbor_f32Data movementNearest-neighbor resize of a float32 NHWC tensor.
arm_resize_nearest_neighbor_f32_get_buffer_sizeData movementScratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16().
arm_resize_nearest_neighbor_s16Data movementNearest neighbor resize function for s16 data.
arm_resize_nearest_neighbor_s8Data movementNearest neighbor resize function for s8 data.
arm_reverse_sequence_s16Data movementReverse variable-length sequences along a dimension for int16.
arm_reverse_sequence_s8Data movementReverse variable-length sequences along a dimension for int8.
arm_rsqrt_f16ElementwiseElementwise reciprocal square root of a float16 tensor, 1 / sqrt(x).
arm_rsqrt_f32ElementwiseElementwise reciprocal square root, 1 / sqrt(x).
arm_rsqrt_s16_per_opElementwiseINT16 reciprocal square root using a per-operator LUT.
arm_rsqrt_s16_universalElementwiseINT16 reciprocal square root using a shared universal LUT.
arm_scatter_nd_s16Data movementScatter updates into a zero-initialized output tensor for int16.
arm_scatter_nd_s8Data movementScatter updates into a zero-initialized output tensor for int8.
arm_select_v2_s16ElementwiseSELECTV2 with broadcast for int16 tensors.
arm_select_v2_s8ElementwiseSELECTV2 with broadcast for int8 tensors.
arm_softmax_f16Pooling, softmax, quantizationSoftmax using the float-native API signature.
arm_softmax_f32Pooling, softmax, quantizationSoftmax using the float-native API signature.
arm_softmax_s16Pooling, softmax, quantizationS16 softmax function.
arm_softmax_s8Pooling, softmax, quantizationS8 softmax function.
arm_softmax_s8_s16Pooling, softmax, quantizationS8 to s16 softmax function.
arm_softmax_u8Pooling, softmax, quantizationU8 softmax function.
arm_space_to_batch_nd_s16Data movementSpace to Batch ND function for s16 data type.
arm_space_to_batch_nd_s8Data movementSpace to Batch ND function for s8 data type.
arm_space_to_depth_s16Data movementSpace to Depth function for s16 data type.
arm_space_to_depth_s8Data movementSpace to Depth function for s8 data type.
arm_split_f16Data movementSplit a float32 tensor of any rank into several tensors along one axis.
arm_split_f32Data movementSplit a float32 tensor of any rank into several tensors along one axis.
arm_split_s16Data movementint16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis
arm_split_s8Data movementint8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis
arm_sqrt_s16Elementwises16 elementwise square root using piecewise LUT with linear interpolation
arm_sqrt_s16_tablefreeElementwises16 elementwise square root without a lookup table
arm_sqrt_s8Elementwises8 elementwise square root
arm_squared_difference_s16Elementwises16 elementwise squared difference of two tensors with support for broadcasting.
arm_squared_difference_s8Elementwises8 elementwise squared difference of two tensors with support for broadcasting.
arm_squared_difference_scalar_s16Elementwises16 elementwise squared difference of scalar and vector.
arm_squared_difference_scalar_s8Elementwises8 elementwise squared difference of scalar and vector.
arm_strided_slice_f16Data movementStrided slice for float32 data (pure copy, TensorFlow Lite compatible).
arm_strided_slice_f32Data movementStrided slice for float32 data (pure copy, TensorFlow Lite compatible).
arm_strided_slice_s16Data movementStrided slice function for int16 data.
arm_strided_slice_s32Data movementStrided slice function for int32 data.
arm_strided_slice_s8Data movementStrided slice function for int8 data.
arm_sub_s16Elementwises16 elementwise subtraction of two tensors with support for broadcasting.
arm_sub_s8Elementwises8 elementwise subtraction of two tensors with support for broadcasting.
arm_sub_scalar_s16Elementwises16 elementwise subtract of scalar and vector (scalar - vector)
arm_sub_scalar_s8Elementwises8 elementwise subtract of scalar and vector (scalar - vector)
arm_svdf_f16SequenceStateful singular value decomposition filter, float16 variant.
arm_svdf_f16_input_ctx_get_buffer_sizeSequenceGet size of the inputctx staging buffer required by armsvdff16().
arm_svdf_f16_output_ctx_get_buffer_sizeSequenceGet size of the outputctx staging buffer required by armsvdff16().
arm_svdf_f32SequenceStateful singular value decomposition filter.
arm_svdf_f32_input_ctx_get_buffer_sizeSequenceGet size of the inputctx staging buffer required by armsvdff32().
arm_svdf_f32_output_ctx_get_buffer_sizeSequenceGet size of the outputctx staging buffer required by armsvdff32().
arm_svdf_s8Sequences8 SVDF function with 8 bit state tensor and 8 bit time weights
arm_svdf_s8_get_buffer_sizeSequenceGet size of the kernel-sum buffer required by armsvdfs8().
arm_svdf_s8_get_buffer_size_dspSequenceGet size of the kernel-sum buffer required by armsvdfs8() for processors with DSP extension.
arm_svdf_s8_get_buffer_size_mveSequenceGet size of the kernel-sum buffer required by armsvdfs8() for Arm(R) Helium Architecture case.
arm_svdf_s8_input_ctx_get_buffer_sizeSequenceGet size of the inputctx staging buffer required by armsvdfs8().
arm_svdf_s8_output_ctx_get_buffer_sizeSequenceGet size of the outputctx staging buffer required by armsvdfs8().
arm_svdf_state_s16_s8Sequences8 SVDF function with 16 bit state tensor and 16 bit time weights
arm_svdf_state_s16_s8_input_ctx_get_buffer_sizeSequenceGet size of the inputctx staging buffer required by armsvdfstates16s8().
arm_svdf_state_s16_s8_output_ctx_get_buffer_sizeSequenceGet size of the outputctx staging buffer required by armsvdfstates16s8().
arm_tanh_s16ActivationTanh activation function for s16.
arm_tile_s16Data movementTile an int16 tensor along each dimension.
arm_tile_s8Data movementTile an int8 tensor along each dimension.
arm_transpose_conv_f16ConvolutionTranspose convolution, dispatch by layout.
arm_transpose_conv_f16_get_buffer_sizeConvolutionGet the temporary buffer size required by transpose convolution.
arm_transpose_conv_f16_get_reverse_conv_buffer_sizeConvolutionGet the reverse-convolution workspace size used by transpose convolution helpers.
arm_transpose_conv_f32ConvolutionTranspose convolution, dispatch by layout.
arm_transpose_conv_f32_get_buffer_sizeConvolutionGet the temporary buffer size required by transpose convolution.
arm_transpose_conv_f32_get_reverse_conv_buffer_sizeConvolutionGet the reverse-convolution workspace size used by transpose convolution helpers.
arm_transpose_conv_nhwc_f16ConvolutionTranspose convolution, NHWC layout.
arm_transpose_conv_nhwc_f32ConvolutionTranspose convolution, NHWC layout.
arm_transpose_conv_s8ConvolutionBasic s8 transpose convolution function.
arm_transpose_conv_s8_get_buffer_sizeConvolutionGet the required buffer size for ctx in s8 transpose conv function.
arm_transpose_conv_s8_get_buffer_size_mveConvolutionGet size of additional buffer required by armtransposeconvs8() for Arm(R) Helium Architecture case.
arm_transpose_conv_s8_get_reverse_conv_buffer_sizeConvolutionGet the required buffer size for outputctx in s8 transpose conv function.
arm_transpose_conv_wrapper_f16ConvolutionTranspose convolution wrapper using the CMSIS-NN baseline path.
arm_transpose_conv_wrapper_f32ConvolutionTranspose convolution wrapper using the CMSIS-NN baseline path.
arm_transpose_conv_wrapper_s8ConvolutionWrapper to select optimal transposed convolution algorithm depending on parameters.
arm_transpose_f16Data movementTranspose a floating-point tensor.
arm_transpose_f32Data movementTranspose a floating-point tensor.
arm_transpose_s16Data movementBasic s16 transpose function.
arm_transpose_s8Data movementBasic transpose function.
arm_unpack_f16Data movementUnstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK).
arm_unpack_f32Data movementUnstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK).
arm_vector_sum_s8Reduction and comparisonCalculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s32 biasdata.
arm_vector_sum_s8_s64Reduction and comparisonCalculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s64 biasdata.
arm_where_s16Reduction and comparisonWHERE operator: return coordinates of non-zero elements in condition (int16).
arm_where_s8Reduction and comparisonWHERE operator: return coordinates of non-zero elements in condition.