Skip to content
heliaCORE
API reference
HELIA HUB

API

Public kernel functions declared in arm_nnfunctions*.h: 426 across 2 headers. Start with an operator family, then follow a function through to its parameters, return values and source.

Machine-readable model · llms.txt · llms-full.txt

Conv2D, depthwise, transpose convolution, wrappers, and buffer helpers. 126 functions.

Function Summary Module
arm_convolve_1_x_n_f16 1xN convolution, dispatch by layout. Convolution Functions
arm_convolve_1_x_n_f16_acc16 1xN convolution, dispatch by layout. Convolution Functions
arm_convolve_1_x_n_f16_get_buffer_size Get the buffer size required by 1xN convolution. Convolution Functions
arm_convolve_1_x_n_f32 1xN convolution, dispatch by layout. Convolution Functions
arm_convolve_1_x_n_f32_get_buffer_size Get the buffer size required by 1xN convolution. Convolution Functions
arm_convolve_1_x_n_nhwc_f16 1xN convolution, NHWC layout. Convolution Functions
arm_convolve_1_x_n_nhwc_f16_acc16 1xN convolution, NHWC layout. Convolution Functions
arm_convolve_1_x_n_nhwc_f32 1xN convolution, NHWC layout. Convolution Functions
arm_convolve_1_x_n_s4 1xn convolution for s4 weights arm_nnfunctions.h
arm_convolve_1_x_n_s4_get_buffer_size Get the required additional buffer size for 1xn convolution. arm_nnfunctions.h
arm_convolve_1_x_n_s8 1xn convolution arm_nnfunctions.h
arm_convolve_1_x_n_s8_get_buffer_size Get the required additional buffer size for 1xn convolution. arm_nnfunctions.h
arm_convolve_1x1_f16 1x1 convolution, dispatch by layout. Convolution Functions
arm_convolve_1x1_f16_acc16 1x1 convolution, dispatch by layout. Convolution Functions
arm_convolve_1x1_f16_get_buffer_size Get the buffer size required by 1x1 convolution. Convolution Functions
arm_convolve_1x1_f32 1x1 convolution, dispatch by layout. Convolution Functions
arm_convolve_1x1_f32_get_buffer_size Get the buffer size required by 1x1 convolution. Convolution Functions
arm_convolve_1x1_nhwc_f16 1x1 convolution, NHWC layout. Convolution Functions
arm_convolve_1x1_nhwc_f16_acc16 1x1 convolution, NHWC layout. Convolution Functions
arm_convolve_1x1_nhwc_f32 1x1 convolution, NHWC layout. Convolution Functions
arm_convolve_1x1_out_s8 Optimised convolution for 1x1 output images (shape of BX1x1xCOUT) for 8x8 computations. arm_nnfunctions.h
arm_convolve_1x1_out_s8_get_buffer_size Get the required scratch buffer size for armconvolve1x1outs8(). arm_nnfunctions.h
arm_convolve_1x1_s16_ns_np_nd Pointwise s16 convolution function: no stride, no padding, no dilation. arm_nnfunctions.h
arm_convolve_1x1_s4 s4 version for 1x1 convolution with support for non-unity stride values arm_nnfunctions.h
arm_convolve_1x1_s4_fast Fast s4 version for 1x1 convolution (non-square shape). arm_nnfunctions.h
arm_convolve_1x1_s4_fast_get_buffer_size Get the required buffer size for armconvolve1x1s4fast. arm_nnfunctions.h
arm_convolve_1x1_s8 s8 version for 1x1 convolution with support for non-unity stride values arm_nnfunctions.h
arm_convolve_1x1_s8_fast Fast s8 version for 1x1 convolution (non-square shape). arm_nnfunctions.h
arm_convolve_1x1_s8_fast_get_buffer_size Get the required buffer size for armconvolve1x1s8fast. arm_nnfunctions.h
arm_convolve_even_s4 Basic s4 convolution function with a requirement of even number of kernels. arm_nnfunctions.h
arm_convolve_even_s4_get_buffer_size Get the required buffer size for armconvolveevens4. arm_nnfunctions.h
arm_convolve_f16 Convolution, dispatch by layout. Convolution Functions
arm_convolve_f16_acc16 Convolution, dispatch by layout. Convolution Functions
arm_convolve_f16_get_buffer_size Get the temporary buffer size required by convolution. Convolution Functions
arm_convolve_f32 Convolution, dispatch by layout. Convolution Functions
arm_convolve_f32_get_buffer_size Get the temporary buffer size required by convolution. Convolution Functions
arm_convolve_nhwc_f16 Convolution, NHWC layout. Convolution Functions
arm_convolve_nhwc_f16_acc16 Convolution, NHWC layout. Convolution Functions
arm_convolve_nhwc_f32 Convolution, NHWC layout. Convolution Functions
arm_convolve_s16 Basic s16 convolution function. arm_nnfunctions.h
arm_convolve_s16_fast_small_kernel armconvolves16fastsmallkernel function. arm_nnfunctions.h
arm_convolve_s16_get_buffer_size Get the required buffer size for s16 convolution function. arm_nnfunctions.h
arm_convolve_s16_group_ch_mult_1 s16 grouped convolution optimized for the case where filterdims->c == 1 and inputch == outputch (channel multiplier = 1). arm_nnfunctions.h
arm_convolve_s4 Basic s4 convolution function. arm_nnfunctions.h
arm_convolve_s4_get_buffer_size Get the required buffer size for s4 convolution function. arm_nnfunctions.h
arm_convolve_s8 Basic s8 convolution function. arm_nnfunctions.h
arm_convolve_s8_3x3_c16_s1 s8 3x3 convolution over 16 input channels with unit stride. arm_nnfunctions.h
arm_convolve_s8_get_buffer_size Get the required buffer size for s8 convolution function. arm_nnfunctions.h
arm_convolve_s8_get_buffer_size_mve Get the required buffer size for armconvolves8 for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_convolve_s8_get_weights_sum_size Get the required buffer size for s8 convolution and depthwise convolution weight sum. arm_nnfunctions.h
arm_convolve_s8_small_cin s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model. arm_nnfunctions.h
arm_convolve_weight_sum Pre-computes per-output-channel weight sums for a standard convolution. arm_nnfunctions.h
arm_convolve_wrapper_f16 Convolution wrapper using the CMSIS-NN baseline path. Convolution Functions
arm_convolve_wrapper_f16_acc16 Convolution wrapper using the CMSIS-NN baseline path. Convolution Functions
arm_convolve_wrapper_f16_get_buffer_size Get the buffer size required by the convolution wrapper. Convolution Functions
arm_convolve_wrapper_f32 Convolution wrapper using the CMSIS-NN baseline path. Convolution Functions
arm_convolve_wrapper_f32_get_buffer_size Get the buffer size required by the convolution wrapper. Convolution Functions
arm_convolve_wrapper_s16 s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution. arm_nnfunctions.h
arm_convolve_wrapper_s16_get_buffer_size Get the required buffer size for armconvolvewrappers16. arm_nnfunctions.h
arm_convolve_wrapper_s16_get_buffer_size_dsp Get the required buffer size for armconvolvewrappers16 for for processors with DSP extension. arm_nnfunctions.h
arm_convolve_wrapper_s16_get_buffer_size_mve Get the required buffer size for armconvolvewrappers16 for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_convolve_wrapper_s4 s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution. arm_nnfunctions.h
arm_convolve_wrapper_s4_get_buffer_size Get the required buffer size for armconvolvewrappers4. arm_nnfunctions.h
arm_convolve_wrapper_s4_get_buffer_size_dsp Get the required buffer size for armconvolvewrappers4 for processors with DSP extension. arm_nnfunctions.h
arm_convolve_wrapper_s4_get_buffer_size_mve Get the required buffer size for armconvolvewrappers4 for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_convolve_wrapper_s8 s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution. arm_nnfunctions.h
arm_convolve_wrapper_s8_get_buffer_size Get the required buffer size for armconvolvewrappers8. arm_nnfunctions.h
arm_convolve_wrapper_s8_get_buffer_size_dsp Get the required buffer size for armconvolvewrappers8 for processors with DSP extension. arm_nnfunctions.h
arm_convolve_wrapper_s8_get_buffer_size_mve Get the required buffer size for armconvolvewrappers8 for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_depthwise_conv_3x3_s8 Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below). arm_nnfunctions.h
arm_depthwise_conv_f16 Depthwise convolution, dispatch by layout. Convolution Functions
arm_depthwise_conv_f16_acc16 Depthwise convolution, dispatch by layout. Convolution Functions
arm_depthwise_conv_f16_get_buffer_size Get the temporary buffer size required by depthwise convolution. Convolution Functions
arm_depthwise_conv_f32 Depthwise convolution, dispatch by layout. Convolution Functions
arm_depthwise_conv_f32_get_buffer_size Get the temporary buffer size required by depthwise convolution. Convolution Functions
arm_depthwise_conv_fast_s16 Optimized s16 depthwise convolution function with constraint that inchannel equals outchannel. arm_nnfunctions.h
arm_depthwise_conv_fast_s16_get_buffer_size Get the required buffer size for optimized s16 depthwise convolution function with constraint that inchannel equals outchannel. arm_nnfunctions.h
arm_depthwise_conv_s16 Basic s16 depthwise convolution function that doesn’t have any constraints on the input dimensions. arm_nnfunctions.h
arm_depthwise_conv_s4 Basic s4 depthwise convolution function that doesn’t have any constraints on the input dimensions. arm_nnfunctions.h
arm_depthwise_conv_s4_opt Optimized s4 depthwise convolution function with constraint that inchannel equals outchannel. arm_nnfunctions.h
arm_depthwise_conv_s4_opt_get_buffer_size Get the required buffer size for optimized s4 depthwise convolution function with constraint that inchannel equals outchannel. arm_nnfunctions.h
arm_depthwise_conv_s8 Basic s8 depthwise convolution function that doesn’t have any constraints on the input dimensions. arm_nnfunctions.h
arm_depthwise_conv_s8_opt Optimized s8 depthwise convolution function with constraint that inchannel equals outchannel. arm_nnfunctions.h
arm_depthwise_conv_s8_opt_3x3 s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of armdepthwiseconvs8opt(), for layers in its gate. arm_nnfunctions.h
arm_depthwise_conv_s8_opt_3x3_c64_s1 armdepthwiseconvs8opt3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps. arm_nnfunctions.h
arm_depthwise_conv_s8_opt_3x3_get_buffer_size Get the scratch size in bytes of armdepthwiseconvs8opt3x3() and armdepthwiseconvs8opt3x3c64s1(). arm_nnfunctions.h
arm_depthwise_conv_s8_opt_channelwise The channel-vectorized path of armdepthwiseconvs8opt() on its own, without the planar attempt. arm_nnfunctions.h
arm_depthwise_conv_s8_opt_get_buffer_size Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel. arm_nnfunctions.h
arm_depthwise_conv_s8_opt_planar The planar path of armdepthwiseconvs8opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane. arm_nnfunctions.h
arm_depthwise_conv_s8_opt_planar_supported Whether armdepthwiseconvs8opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels. arm_nnfunctions.h
arm_depthwise_conv_wrapper_f16 Depthwise convolution wrapper using the CMSIS-NN baseline path. Convolution Functions
arm_depthwise_conv_wrapper_f16_acc16 Depthwise convolution wrapper using the CMSIS-NN baseline path. Convolution Functions
arm_depthwise_conv_wrapper_f16_get_buffer_size Get the buffer size required by the depthwise convolution wrapper. Convolution Functions
arm_depthwise_conv_wrapper_f32 Depthwise convolution wrapper using the CMSIS-NN baseline path. Convolution Functions
arm_depthwise_conv_wrapper_f32_get_buffer_size Get the buffer size required by the depthwise convolution wrapper. Convolution Functions
arm_depthwise_conv_wrapper_s16 Wrapper function to pick the right optimized s16 depthwise convolution function. arm_nnfunctions.h
arm_depthwise_conv_wrapper_s16_get_buffer_size Get size of additional buffer required by armdepthwiseconvwrappers16(). arm_nnfunctions.h
arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp Get size of additional buffer required by armdepthwiseconvwrappers16() for processors with DSP extension. arm_nnfunctions.h
arm_depthwise_conv_wrapper_s16_get_buffer_size_mve Get size of additional buffer required by armdepthwiseconvwrappers16() for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_depthwise_conv_wrapper_s4 Wrapper function to pick the right optimized s4 depthwise convolution function. arm_nnfunctions.h
arm_depthwise_conv_wrapper_s4_get_buffer_size Get size of additional buffer required by armdepthwiseconvwrappers4(). arm_nnfunctions.h
arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp Get size of additional buffer required by armdepthwiseconvwrappers4() for processors with DSP extension. arm_nnfunctions.h
arm_depthwise_conv_wrapper_s4_get_buffer_size_mve Get size of additional buffer required by armdepthwiseconvwrappers4() for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_depthwise_conv_wrapper_s8 Wrapper function to pick the right optimized s8 depthwise convolution function. arm_nnfunctions.h
arm_depthwise_conv_wrapper_s8_get_buffer_size Get size of additional buffer required by armdepthwiseconvwrappers8(). arm_nnfunctions.h
arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp Get size of additional buffer required by armdepthwiseconvwrappers8() for processors with DSP extension. arm_nnfunctions.h
arm_depthwise_conv_wrapper_s8_get_buffer_size_mve Get size of additional buffer required by armdepthwiseconvwrappers8() for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_depthwise_convolve_weight_sum Pre-computes per-channel weight sums for a depthwise convolution. arm_nnfunctions.h
arm_depthwise_nhwc_conv_f16 Depthwise convolution, NHWC layout. Convolution Functions
arm_depthwise_nhwc_conv_f16_acc16 Depthwise convolution, NHWC layout. Convolution Functions
arm_depthwise_nhwc_conv_f32 Depthwise convolution, NHWC layout. Convolution Functions
arm_transpose_conv_f16 Transpose convolution, dispatch by layout. Convolution Functions
arm_transpose_conv_f16_get_buffer_size Get the temporary buffer size required by transpose convolution. Convolution Functions
arm_transpose_conv_f16_get_reverse_conv_buffer_size Get the reverse-convolution workspace size used by transpose convolution helpers. Convolution Functions
arm_transpose_conv_f32 Transpose convolution, dispatch by layout. Convolution Functions
arm_transpose_conv_f32_get_buffer_size Get the temporary buffer size required by transpose convolution. Convolution Functions
arm_transpose_conv_f32_get_reverse_conv_buffer_size Get the reverse-convolution workspace size used by transpose convolution helpers. Convolution Functions
arm_transpose_conv_nhwc_f16 Transpose convolution, NHWC layout. Convolution Functions
arm_transpose_conv_nhwc_f32 Transpose convolution, NHWC layout. Convolution Functions
arm_transpose_conv_s8 Basic s8 transpose convolution function. arm_nnfunctions.h
arm_transpose_conv_s8_get_buffer_size Get the required buffer size for ctx in s8 transpose conv function. arm_nnfunctions.h
arm_transpose_conv_s8_get_buffer_size_mve Get size of additional buffer required by armtransposeconvs8() for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_transpose_conv_s8_get_reverse_conv_buffer_size Get the required buffer size for outputctx in s8 transpose conv function. arm_nnfunctions.h
arm_transpose_conv_wrapper_f16 Transpose convolution wrapper using the CMSIS-NN baseline path. Convolution Functions
arm_transpose_conv_wrapper_f32 Transpose convolution wrapper using the CMSIS-NN baseline path. Convolution Functions
arm_transpose_conv_wrapper_s8 Wrapper to select optimal transposed convolution algorithm depending on parameters. arm_nnfunctions.h

Dense layers, batch matmul paths, and scratch sizing helpers. 33 functions.

Function Summary Module
arm_batch_matmul_f16 Batched matrix multiplication. Fully-connected Layer Functions
arm_batch_matmul_f16_get_buffer_size Get the temporary buffer size required by batched matrix multiplication. Fully-connected Layer Functions
arm_batch_matmul_f32 Batched matrix multiplication. Fully-connected Layer Functions
arm_batch_matmul_f32_get_buffer_size Get the temporary buffer size required by batched matrix multiplication. Fully-connected Layer Functions
arm_batch_matmul_s16 Batch matmul function with 16 bit input and output. arm_nnfunctions.h
arm_batch_matmul_s8 Batch matmul function with 8 bit input and output. arm_nnfunctions.h
arm_batch_matmul_s8_get_buffer_size Get size of the scratch buffer required by armbatchmatmuls8(). arm_nnfunctions.h
arm_batch_matmul_s8_get_buffer_size_dsp Get size of the scratch buffer required by armbatchmatmuls8() for processors with DSP extension. arm_nnfunctions.h
arm_batch_matmul_s8_get_buffer_size_mve Get size of the scratch buffer required by armbatchmatmuls8() for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_fully_connected_f16 Fully connected layer, dispatch by layout. Fully-connected Layer Functions
arm_fully_connected_f16_acc16 Fully connected layer, dispatch by layout. Fully-connected Layer Functions
arm_fully_connected_f16_get_buffer_size Get the temporary buffer size required by the fully connected layer. Fully-connected Layer Functions
arm_fully_connected_f32 Fully connected layer, dispatch by layout. Fully-connected Layer Functions
arm_fully_connected_f32_get_buffer_size Get the temporary buffer size required by the fully connected layer. Fully-connected Layer Functions
arm_fully_connected_nhwc_f16 Fully connected layer, NHWC layout. Fully-connected Layer Functions
arm_fully_connected_nhwc_f16_acc16 Fully connected layer, NHWC layout. Fully-connected Layer Functions
arm_fully_connected_nhwc_f32 Fully connected layer, NHWC layout. Fully-connected Layer Functions
arm_fully_connected_per_channel_s16 Basic s16 Fully Connected function using per channel quantization. arm_nnfunctions.h
arm_fully_connected_per_channel_s16_get_buffer_size Get size of additional buffer required by armfullyconnectedperchannels16(). arm_nnfunctions.h
arm_fully_connected_per_channel_s16_get_buffer_size_dsp Get size of additional buffer required by armfullyconnectedperchannels16() for processors with DSP extension. arm_nnfunctions.h
arm_fully_connected_per_channel_s16_get_buffer_size_mve Get size of additional buffer required by armfullyconnectedperchannels16() for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_fully_connected_per_channel_s8 Basic s8 Fully Connected function using per channel quantization. arm_nnfunctions.h
arm_fully_connected_s16 Basic s16 Fully Connected function. arm_nnfunctions.h
arm_fully_connected_s16_get_buffer_size Get size of additional buffer required by armfullyconnecteds16(). arm_nnfunctions.h
arm_fully_connected_s16_get_buffer_size_dsp Get size of additional buffer required by armfullyconnecteds16() for processors with DSP extension. arm_nnfunctions.h
arm_fully_connected_s16_get_buffer_size_mve Get size of additional buffer required by armfullyconnecteds16() for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_fully_connected_s4 Basic s4 Fully Connected function. arm_nnfunctions.h
arm_fully_connected_s8 Basic s8 Fully Connected function. arm_nnfunctions.h
arm_fully_connected_s8_get_buffer_size Get size of additional buffer required by armfullyconnecteds8(). arm_nnfunctions.h
arm_fully_connected_s8_get_buffer_size_dsp Get size of additional buffer required by armfullyconnecteds8() for processors with DSP extension. arm_nnfunctions.h
arm_fully_connected_s8_get_buffer_size_mve Get size of additional buffer required by armfullyconnecteds8() for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_fully_connected_wrapper_s16 s16 Fully Connected layer wrapper function arm_nnfunctions.h
arm_fully_connected_wrapper_s8 s8 Fully Connected layer wrapper function arm_nnfunctions.h

Add, sub, mul, square difference, min/max, batch norm, select, and arithmetic glue. 65 functions.

Function Summary Module
arm_abs_s16 s16 elementwise absolute value arm_nnfunctions.h
arm_abs_s8 s8 elementwise absolute value arm_nnfunctions.h
arm_add_s16 s16 elementwise add of two tensors with support for broadcasting. arm_nnfunctions.h
arm_add_s8 s8 elementwise add of two tensors with support for broadcasting. arm_nnfunctions.h
arm_add_scalar_s16 s16 elementwise add of scalar and vector arm_nnfunctions.h
arm_add_scalar_s8 s8 elementwise add of scalar and vector arm_nnfunctions.h
arm_batch_norm_f16 Apply batch normalization. NNSupport
arm_batch_norm_f32 Apply batch normalization. NNSupport
arm_elementwise_add_broadcast_f16 Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp. Elementwise Functions
arm_elementwise_add_broadcast_f32 Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp. Elementwise Functions
arm_elementwise_add_f16 Elementwise add with optional output clamp. Elementwise Functions
arm_elementwise_add_f32 Elementwise add with optional output clamp. Elementwise Functions
arm_elementwise_add_fp16 Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate armelementwiseaddf16(). Elementwise Functions
arm_elementwise_add_s16 s16 elementwise add of two vectors arm_nnfunctions.h
arm_elementwise_add_s8 s8 elementwise add of two vectors arm_nnfunctions.h
arm_elementwise_mul_broadcast_f16 Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp. Elementwise Functions
arm_elementwise_mul_broadcast_f32 Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp. Elementwise Functions
arm_elementwise_mul_f16 Elementwise multiply with optional output clamp. Elementwise Functions
arm_elementwise_mul_f32 Elementwise multiply with optional output clamp. Elementwise Functions
arm_elementwise_mul_s16 s16 elementwise multiplication arm_nnfunctions.h
arm_elementwise_mul_s8 s8 elementwise multiplication arm_nnfunctions.h
arm_elementwise_prelu_s16 Elementwise S16 PReLU activation function. arm_nnfunctions.h
arm_elementwise_prelu_s8 Elementwise S8 PReLU activation function. arm_nnfunctions.h
arm_elementwise_squared_difference_f16 Elementwise squared difference of two float16 vectors. Elementwise Functions
arm_elementwise_squared_difference_s16 s16 elementwise squared difference of two vectors. arm_nnfunctions.h
arm_elementwise_squared_difference_s8 s8 elementwise squared difference of two vectors. arm_nnfunctions.h
arm_elementwise_sub_broadcast_f16 Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp. Elementwise Functions
arm_elementwise_sub_broadcast_f32 Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp. Elementwise Functions
arm_elementwise_sub_f16 Elementwise subtract with optional output clamp. Elementwise Functions
arm_elementwise_sub_f32 Elementwise subtract with optional output clamp. Elementwise Functions
arm_elementwise_sub_s16 s16 elementwise subtract of two vectors arm_nnfunctions.h
arm_elementwise_sub_s8 s8 elementwise subtract of two vectors arm_nnfunctions.h
arm_maximum_f16 Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. Elementwise Functions
arm_maximum_f32 Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. Elementwise Functions
arm_maximum_s16 s16 elementwise maximum w/ support for broadcasting and scalar inputs. arm_nnfunctions.h
arm_maximum_s8 s8 elementwise maximum w/ support for broadcasting and scalar inputs. arm_nnfunctions.h
arm_minimum_f16 Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. Elementwise Functions
arm_minimum_f32 Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. Elementwise Functions
arm_minimum_s16 s16 elementwise minimum w/ support for broadcasting and scalar inputs. arm_nnfunctions.h
arm_minimum_s8 s8 elementwise minimum w/ support for broadcasting and scalar inputs. arm_nnfunctions.h
arm_mul_s16 s16 elementwise multiplication of two tensors with support for broadcasting. arm_nnfunctions.h
arm_mul_s8 s8 elementwise multiplication of two tensors with support for broadcasting. arm_nnfunctions.h
arm_mul_scalar_s16 s16 elementwise multiplication of scalar and vector arm_nnfunctions.h
arm_mul_scalar_s8 s8 elementwise multiplication of scalar and vector arm_nnfunctions.h
arm_nn_abs_f16 Elementwise absolute value. Elementwise Functions
arm_nn_abs_f32 Elementwise absolute value. Elementwise Functions
arm_nn_sqrt_f16 Elementwise square root of a float16 tensor. Elementwise Functions
arm_nn_sqrt_f32 Elementwise square root. Elementwise Functions
arm_rsqrt_f16 Elementwise reciprocal square root of a float16 tensor, 1 / sqrt(x). Elementwise Functions
arm_rsqrt_f32 Elementwise reciprocal square root, 1 / sqrt(x). Elementwise Functions
arm_rsqrt_s16_per_op INT16 reciprocal square root using a per-operator LUT. arm_nnfunctions.h
arm_rsqrt_s16_universal INT16 reciprocal square root using a shared universal LUT. arm_nnfunctions.h
arm_select_v2_s16 SELECTV2 with broadcast for int16 tensors. arm_nnfunctions.h
arm_select_v2_s8 SELECTV2 with broadcast for int8 tensors. arm_nnfunctions.h
arm_sqrt_s16 s16 elementwise square root using piecewise LUT with linear interpolation arm_nnfunctions.h
arm_sqrt_s16_tablefree s16 elementwise square root without a lookup table arm_nnfunctions.h
arm_sqrt_s8 s8 elementwise square root arm_nnfunctions.h
arm_squared_difference_s16 s16 elementwise squared difference of two tensors with support for broadcasting. arm_nnfunctions.h
arm_squared_difference_s8 s8 elementwise squared difference of two tensors with support for broadcasting. arm_nnfunctions.h
arm_squared_difference_scalar_s16 s16 elementwise squared difference of scalar and vector. arm_nnfunctions.h
arm_squared_difference_scalar_s8 s8 elementwise squared difference of scalar and vector. arm_nnfunctions.h
arm_sub_s16 s16 elementwise subtraction of two tensors with support for broadcasting. arm_nnfunctions.h
arm_sub_s8 s8 elementwise subtraction of two tensors with support for broadcasting. arm_nnfunctions.h
arm_sub_scalar_s16 s16 elementwise subtract of scalar and vector (scalar - vector) arm_nnfunctions.h
arm_sub_scalar_s8 s8 elementwise subtract of scalar and vector (scalar - vector) arm_nnfunctions.h

Argmin/argmax, min/max reductions, comparisons, means, vector sums, and where. 40 functions.

Function Summary Module
arm_argmax_f16 Returns the first maximum’s axis-relative INT32 index for a f16 tensor. Reduction Functions
arm_argmax_f32 Returns the first maximum’s axis-relative INT32 index for a f32 tensor. Reduction Functions
arm_argmax_s16 Compute ArgMax indices of an s16 tensor along a specific axis. arm_nnfunctions.h
arm_argmax_s8 Compute ArgMax indices of an s8 tensor along a specific axis. arm_nnfunctions.h
arm_argmin_f16 Returns the first minimum’s axis-relative INT32 index for a f16 tensor. Reduction Functions
arm_argmin_f32 Returns the first minimum’s axis-relative INT32 index for a f32 tensor. Reduction Functions
arm_argmin_s16 Compute ArgMin indices of an s16 tensor along a specific axis. arm_nnfunctions.h
arm_argmin_s8 Compute ArgMin indices of an s8 tensor along a specific axis. arm_nnfunctions.h
arm_comparison_s16 s16 elementwise comparison with support for broadcasting. arm_nnfunctions.h
arm_comparison_s8 s8 elementwise comparison with support for broadcasting. arm_nnfunctions.h
arm_equal_s16 s16 elementwise equality comparison with support for broadcasting. arm_nnfunctions.h
arm_equal_s8 s8 elementwise equality comparison with support for broadcasting. arm_nnfunctions.h
arm_greater_equal_s16 s16 elementwise greater-or-equal comparison with support for broadcasting. arm_nnfunctions.h
arm_greater_equal_s8 s8 elementwise greater-or-equal comparison with support for broadcasting. arm_nnfunctions.h
arm_greater_s16 s16 elementwise greater-than comparison with support for broadcasting. arm_nnfunctions.h
arm_greater_s8 s8 elementwise greater-than comparison with support for broadcasting. arm_nnfunctions.h
arm_less_equal_s16 s16 elementwise less-or-equal comparison with support for broadcasting. arm_nnfunctions.h
arm_less_equal_s8 s8 elementwise less-or-equal comparison with support for broadcasting. arm_nnfunctions.h
arm_less_s16 s16 elementwise less-than comparison with support for broadcasting. arm_nnfunctions.h
arm_less_s8 s8 elementwise less-than comparison with support for broadcasting. arm_nnfunctions.h
arm_mean_s16 Computes the mean of the input tensor along the specified axis. arm_nnfunctions.h
arm_mean_s8 Computes the mean of the input tensor along the specified axis. arm_nnfunctions.h
arm_nn_mean_f16 Computes the mean of a float16 tensor along the specified axes. Reduction Functions
arm_nn_mean_f32 Computes the mean of a float32 tensor along the specified axes. Reduction Functions
arm_not_equal_s16 s16 elementwise inequality comparison with support for broadcasting. arm_nnfunctions.h
arm_not_equal_s8 s8 elementwise inequality comparison with support for broadcasting. arm_nnfunctions.h
arm_reduce_max_f16 Reduces a f16 NHWC tensor to its maximum along a binary axis mask. Reduction Functions
arm_reduce_max_f32 Reduces a f32 NHWC tensor to its maximum along a binary axis mask. Reduction Functions
arm_reduce_max_s16 Computes the max of the input tensor along the specified axis. arm_nnfunctions.h
arm_reduce_max_s8 Computes the max of the input tensor along the specified axis. arm_nnfunctions.h
arm_reduce_min_f16 Reduces a f16 NHWC tensor to its minimum along a binary axis mask. Reduction Functions
arm_reduce_min_f32 Reduces a f32 NHWC tensor to its minimum along a binary axis mask. Reduction Functions
arm_reduce_min_s16 Computes the min of the input tensor along the specified axis. arm_nnfunctions.h
arm_reduce_min_s8 Computes the min of the input tensor along the specified axis. arm_nnfunctions.h
arm_reduce_sum_f16 Computes the sum of the input tensor along the specified axes. Reduction Functions
arm_reduce_sum_f32 Computes the sum of the input tensor along the specified axes. Reduction Functions
arm_vector_sum_s8 Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s32 biasdata. arm_nnfunctions.h
arm_vector_sum_s8_s64 Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s64 biasdata. arm_nnfunctions.h
arm_where_s16 WHERE operator: return coordinates of non-zero elements in condition (int16). arm_nnfunctions.h
arm_where_s8 WHERE operator: return coordinates of non-zero elements in condition. arm_nnfunctions.h

ReLU, LeakyReLU, PReLU, Hard-Swish, Logistic, Tanh, and clamp. 27 functions.

Function Summary Module
arm_clamp_s16 S16 clamp function. arm_nnfunctions.h
arm_clamp_s8 S8 clamp function. arm_nnfunctions.h
arm_hard_swish_compat_s8 S8 Hard-Swish activation function (compatibility version). arm_nnfunctions.h
arm_hard_swish_f16 Hard swish activation for float16 data. Activation Functions
arm_hard_swish_f32 Hard swish activation for float32 data. Activation Functions
arm_hard_swish_precise_s16 S16 Hard-Swish activation function (precise version). arm_nnfunctions.h
arm_hard_swish_precise_s8 S8 Hard-Swish activation function (precise version). arm_nnfunctions.h
arm_leaky_relu_s16 S16 Leaky ReLU activation function. arm_nnfunctions.h
arm_leaky_relu_s8 S8 Leaky ReLU activation function. arm_nnfunctions.h
arm_logistic_s16 Logistic activation function for s16. arm_nnfunctions.h
arm_nn_activation_f16 Elementwise activation. Activation Functions
arm_nn_activation_f32 Elementwise activation. Activation Functions
arm_nn_activation_s16 s16 neural network activation function using direct table look-up arm_nnfunctions.h
arm_prelu_f16 Parametric ReLU for float32 data. Activation Functions
arm_prelu_f32 Parametric ReLU for float32 data. Activation Functions
arm_prelu_s16 S16 PReLU activation function. arm_nnfunctions.h
arm_prelu_s8 S8 PReLU activation function. arm_nnfunctions.h
arm_prelu_scalar_s16 Scalar S16 PReLU activation function. arm_nnfunctions.h
arm_prelu_scalar_s8 Scalar S8 PReLU activation function. arm_nnfunctions.h
arm_relu_generic_s16 S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc. arm_nnfunctions.h
arm_relu_generic_s8 S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc. arm_nnfunctions.h
arm_relu_q15 Q15 RELU function. arm_nnfunctions.h
arm_relu_q7 Q7 RELU function. arm_nnfunctions.h
arm_relu_s16 S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767. arm_nnfunctions.h
arm_relu_s8 S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127. arm_nnfunctions.h
arm_relu6_q7 Q7 RELU6 function. arm_nnfunctions.h
arm_tanh_s16 Tanh activation function for s16. arm_nnfunctions.h

Pad, reshape, transpose, concatenate, gather, resize, strided slice, tile, broadcast, scatter, mirror pad, and sequence/slice-update utilities. 77 functions.

Function Summary Module
arm_batch_to_space_nd_s16 Batch to Space ND function for s16 data type. arm_nnfunctions.h
arm_batch_to_space_nd_s8 Batch to Space ND function for s8 data type. arm_nnfunctions.h
arm_broadcast_to_s16 Broadcast an int16 tensor to a target shape. arm_nnfunctions.h
arm_broadcast_to_s8 Broadcast an int8 tensor to a target shape. arm_nnfunctions.h
arm_concatenation_f16 Concatenate float32 tensors of any rank along one axis. NNSupport
arm_concatenation_f16_w Concatenate tensors along the W axis. NNSupport
arm_concatenation_f16_x Concatenate tensors along the X axis. NNSupport
arm_concatenation_f16_y Concatenate tensors along the Y axis. NNSupport
arm_concatenation_f16_z Concatenate tensors along the Z axis. NNSupport
arm_concatenation_f32 Concatenate float32 tensors of any rank along one axis. NNSupport
arm_concatenation_f32_w Concatenate tensors along the W axis. NNSupport
arm_concatenation_f32_x Concatenate tensors along the X axis. NNSupport
arm_concatenation_f32_y Concatenate tensors along the Y axis. NNSupport
arm_concatenation_f32_z Concatenate tensors along the Z axis. NNSupport
arm_concatenation_s16 int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis arm_nnfunctions.h
arm_concatenation_s32 int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis arm_nnfunctions.h
arm_concatenation_s8 int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis arm_nnfunctions.h
arm_concatenation_s8_w int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to… arm_nnfunctions.h
arm_concatenation_s8_x int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate. arm_nnfunctions.h
arm_concatenation_s8_y int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate. arm_nnfunctions.h
arm_concatenation_s8_z int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate. arm_nnfunctions.h
arm_depth_to_space_s16 Depth to Space function for s16 data type. arm_nnfunctions.h
arm_depth_to_space_s8 Depth to Space function for s8 data type. arm_nnfunctions.h
arm_dynamic_update_slice_s16 Update a slice of an int16 operand tensor at runtime-determined indices. arm_nnfunctions.h
arm_dynamic_update_slice_s8 Update a slice of an int8 operand tensor at runtime-determined indices. arm_nnfunctions.h
arm_gather_f16 Gather contiguous slices along an axis. Gather Functions:
arm_gather_f32 Gather contiguous slices along an axis. Gather Functions:
arm_gather_nd_f16 Gather contiguous slices using coordinate tuples. Gather Functions:
arm_gather_nd_f32 Gather contiguous slices using coordinate tuples. Gather Functions:
arm_gather_nd_s16 Gathernd slices for int16 tensors. arm_nnfunctions.h
arm_gather_nd_s8 Gathernd slices for int8 tensors. arm_nnfunctions.h
arm_gather_s16 Gather elements along an axis for int16 tensors. arm_nnfunctions.h
arm_gather_s8 Gather elements along an axis for int8 tensors. arm_nnfunctions.h
arm_mirror_pad_s16 Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode). arm_nnfunctions.h
arm_mirror_pad_s8 Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode). arm_nnfunctions.h
arm_nn_fill_f16 Fill a float16 vector with one value; bit copy of value, NaN payload included. Elementwise Functions
arm_nn_fill_f32 Fill a float32 vector with one value. Elementwise Functions
arm_pack_f16 Stack float32 tensors of equal shape along a new axis (TFLite PACK). NNSupport
arm_pack_f32 Stack float32 tensors of equal shape along a new axis (TFLite PACK). NNSupport
arm_pad_f16 Pad a tensor with a constant value. Pad Layer Functions:
arm_pad_f32 Pad a tensor with a constant value. Pad Layer Functions:
arm_pad_s16 Expands the size of the input by adding constant values before and after the data, in all dimensions. arm_nnfunctions.h
arm_pad_s8 Expands the size of the input by adding constant values before and after the data, in all dimensions. arm_nnfunctions.h
arm_reshape_f16 Reshape by copying data without changing element order. NNSupport
arm_reshape_f32 Reshape by copying data without changing element order. NNSupport
arm_reshape_s8 Reshape a s8 vector into another with different shape. arm_nnfunctions.h
arm_resize_nearest_neighbor_f16 Nearest-neighbor resize of a float32 NHWC tensor. Reshape Functions
arm_resize_nearest_neighbor_f16_get_buffer_size Scratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16(). Reshape Functions
arm_resize_nearest_neighbor_f32 Nearest-neighbor resize of a float32 NHWC tensor. Reshape Functions
arm_resize_nearest_neighbor_f32_get_buffer_size Scratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16(). Reshape Functions
arm_resize_nearest_neighbor_s16 Nearest neighbor resize function for s16 data. arm_nnfunctions.h
arm_resize_nearest_neighbor_s8 Nearest neighbor resize function for s8 data. arm_nnfunctions.h
arm_reverse_sequence_s16 Reverse variable-length sequences along a dimension for int16. arm_nnfunctions.h
arm_reverse_sequence_s8 Reverse variable-length sequences along a dimension for int8. arm_nnfunctions.h
arm_scatter_nd_s16 Scatter updates into a zero-initialized output tensor for int16. arm_nnfunctions.h
arm_scatter_nd_s8 Scatter updates into a zero-initialized output tensor for int8. arm_nnfunctions.h
arm_space_to_batch_nd_s16 Space to Batch ND function for s16 data type. arm_nnfunctions.h
arm_space_to_batch_nd_s8 Space to Batch ND function for s8 data type. arm_nnfunctions.h
arm_space_to_depth_s16 Space to Depth function for s16 data type. arm_nnfunctions.h
arm_space_to_depth_s8 Space to Depth function for s8 data type. arm_nnfunctions.h
arm_split_f16 Split a float32 tensor of any rank into several tensors along one axis. Elementwise Functions
arm_split_f32 Split a float32 tensor of any rank into several tensors along one axis. NNSupport
arm_split_s16 int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis arm_nnfunctions.h
arm_split_s8 int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis arm_nnfunctions.h
arm_strided_slice_f16 Strided slice for float32 data (pure copy, TensorFlow Lite compatible). Elementwise Functions
arm_strided_slice_f32 Strided slice for float32 data (pure copy, TensorFlow Lite compatible). Slicing Functions:
arm_strided_slice_s16 Strided slice function for int16 data. arm_nnfunctions.h
arm_strided_slice_s32 Strided slice function for int32 data. arm_nnfunctions.h
arm_strided_slice_s8 Strided slice function for int8 data. arm_nnfunctions.h
arm_tile_s16 Tile an int16 tensor along each dimension. arm_nnfunctions.h
arm_tile_s8 Tile an int8 tensor along each dimension. arm_nnfunctions.h
arm_transpose_f16 Transpose a floating-point tensor. NNSupport
arm_transpose_f32 Transpose a floating-point tensor. NNSupport
arm_transpose_s16 Basic s16 transpose function. arm_nnfunctions.h
arm_transpose_s8 Basic transpose function. arm_nnfunctions.h
arm_unpack_f16 Unstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK). NNSupport
arm_unpack_f32 Unstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK). NNSupport

Classifier tail APIs plus dtype conversion and requantization utilities. 27 functions.

Function Summary Module
arm_avg_pool_f16 Average pooling. Pooling Functions
arm_avg_pool_f32 Average pooling. Pooling Functions
arm_avgpool_s16 s16 average pooling function. arm_nnfunctions.h
arm_avgpool_s16_get_buffer_size Get the required buffer size for S16 average pooling function. arm_nnfunctions.h
arm_avgpool_s16_get_buffer_size_dsp Get the required buffer size for S16 average pooling function for processors with DSP extension. arm_nnfunctions.h
arm_avgpool_s16_get_buffer_size_mve Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_avgpool_s8 s8 average pooling function. arm_nnfunctions.h
arm_avgpool_s8_get_buffer_size Get the required buffer size for S8 average pooling function. arm_nnfunctions.h
arm_avgpool_s8_get_buffer_size_dsp Get the required buffer size for S8 average pooling function for processors with DSP extension. arm_nnfunctions.h
arm_avgpool_s8_get_buffer_size_mve Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_dequantize_f16_f32 Widen a float16 vector to float32. Quantization Functions:
arm_dequantize_s16_f32 Dequantize an int16t array back to floating-point format. arm_nnfunctions.h
arm_dequantize_s8_f32 Dequantize an int8t array back to floating-point format. arm_nnfunctions.h
arm_max_pool_f16 Max pooling. Pooling Functions
arm_max_pool_f32 Max pooling. Pooling Functions
arm_max_pool_s16 s16 max pooling function. arm_nnfunctions.h
arm_max_pool_s8 s8 max pooling function. arm_nnfunctions.h
arm_quantize_f32_s16 Quantize a floating-point array into int16t format. arm_nnfunctions.h
arm_quantize_f32_s8 Quantize a floating-point array into int8t format. arm_nnfunctions.h
arm_requantize_s16_s16 Requantize an int16t array to another int16t range with a different scale. arm_nnfunctions.h
arm_requantize_s8_s8 Requantize an int8t array to another int8t range with a different scale. arm_nnfunctions.h
arm_softmax_f16 Softmax using the float-native API signature. Softmax Functions
arm_softmax_f32 Softmax using the float-native API signature. Softmax Functions
arm_softmax_s16 S16 softmax function. arm_nnfunctions.h
arm_softmax_s8 S8 softmax function. arm_nnfunctions.h
arm_softmax_s8_s16 S8 to s16 softmax function. arm_nnfunctions.h
arm_softmax_u8 U8 softmax function. arm_nnfunctions.h

LSTM, GRU, and SVDF functions for temporal workloads. 31 functions.

Function Summary Module
arm_gru_unidirectional_f16 Unidirectional GRU layer for float16 input, output and state. LSTM Layer Functions
arm_gru_unidirectional_f16_temp1_get_buffer_size Get size of the temp1 scratch buffer required by armgruunidirectionalf16(). LSTM Layer Functions
arm_gru_unidirectional_f32 Unidirectional GRU layer for float32 input, output and state. LSTM Layer Functions
arm_gru_unidirectional_f32_temp1_get_buffer_size Get size of the temp1 scratch buffer required by armgruunidirectionalf32(). LSTM Layer Functions
arm_lstm_unidirectional_f16 Unidirectional LSTM inference. LSTM Layer Functions
arm_lstm_unidirectional_f16_temp1_get_buffer_size Get size of the temp1 scratch buffer required by armlstmunidirectionalf16(). LSTM Layer Functions
arm_lstm_unidirectional_f16_temp2_get_buffer_size Get size of the temp2 scratch buffer required by armlstmunidirectionalf16(). LSTM Layer Functions
arm_lstm_unidirectional_f32 Unidirectional LSTM inference. LSTM Layer Functions
arm_lstm_unidirectional_f32_temp1_get_buffer_size Get size of the temp1 scratch buffer required by armlstmunidirectionalf32(). LSTM Layer Functions
arm_lstm_unidirectional_f32_temp2_get_buffer_size Get size of the temp2 scratch buffer required by armlstmunidirectionalf32(). LSTM Layer Functions
arm_lstm_unidirectional_s16 LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias. arm_nnfunctions.h
arm_lstm_unidirectional_s16_temp1_get_buffer_size Get size of the temp1 scratch buffer required by armlstmunidirectionals16(). arm_nnfunctions.h
arm_lstm_unidirectional_s16_temp2_get_buffer_size Get size of the temp2 scratch buffer required by armlstmunidirectionals16(). arm_nnfunctions.h
arm_lstm_unidirectional_s8 LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias. arm_nnfunctions.h
arm_lstm_unidirectional_s8_temp1_get_buffer_size Get size of the temp1 scratch buffer required by armlstmunidirectionals8(). arm_nnfunctions.h
arm_lstm_unidirectional_s8_temp2_get_buffer_size Get size of the temp2 scratch buffer required by armlstmunidirectionals8(). arm_nnfunctions.h
arm_svdf_f16 Stateful singular value decomposition filter, float16 variant. SVDF Functions
arm_svdf_f16_input_ctx_get_buffer_size Get size of the inputctx staging buffer required by armsvdff16(). SVDF Functions
arm_svdf_f16_output_ctx_get_buffer_size Get size of the outputctx staging buffer required by armsvdff16(). SVDF Functions
arm_svdf_f32 Stateful singular value decomposition filter. SVDF Functions
arm_svdf_f32_input_ctx_get_buffer_size Get size of the inputctx staging buffer required by armsvdff32(). SVDF Functions
arm_svdf_f32_output_ctx_get_buffer_size Get size of the outputctx staging buffer required by armsvdff32(). SVDF Functions
arm_svdf_s8 s8 SVDF function with 8 bit state tensor and 8 bit time weights arm_nnfunctions.h
arm_svdf_s8_get_buffer_size Get size of the kernel-sum buffer required by armsvdfs8(). arm_nnfunctions.h
arm_svdf_s8_get_buffer_size_dsp Get size of the kernel-sum buffer required by armsvdfs8() for processors with DSP extension. arm_nnfunctions.h
arm_svdf_s8_get_buffer_size_mve Get size of the kernel-sum buffer required by armsvdfs8() for Arm(R) Helium Architecture case. arm_nnfunctions.h
arm_svdf_s8_input_ctx_get_buffer_size Get size of the inputctx staging buffer required by armsvdfs8(). arm_nnfunctions.h
arm_svdf_s8_output_ctx_get_buffer_size Get size of the outputctx staging buffer required by armsvdfs8(). arm_nnfunctions.h
arm_svdf_state_s16_s8 s8 SVDF function with 16 bit state tensor and 16 bit time weights arm_nnfunctions.h
arm_svdf_state_s16_s8_input_ctx_get_buffer_size Get size of the inputctx staging buffer required by armsvdfstates16s8(). arm_nnfunctions.h
arm_svdf_state_s16_s8_output_ctx_get_buffer_size Get size of the outputctx staging buffer required by armsvdfstates16s8(). arm_nnfunctions.h

Structs, enums, macros and typedefs stay on the module pages rather than being repeated here. Open heliaCORE for the full tree.