API
Public kernel functions declared in arm_nnfunctions*.h: 426 across 2 headers. Start with an operator family, then follow a function through to its parameters, return values and source.
Machine-readable model · llms.txt · llms-full.txt
compute
Convolution
Conv2D, depthwise, transpose convolution, wrappers, and buffer helpers.
compute
Fully connected
Dense layers, batch matmul paths, and scratch sizing helpers.
graph glue
Elementwise
Add, sub, mul, square difference, min/max, batch norm, select, and arithmetic glue.
selection
Reduction and comparison
Argmin/argmax, min/max reductions, comparisons, means, vector sums, and where.
activation
Activation
ReLU, LeakyReLU, PReLU, Hard-Swish, Logistic, Tanh, and clamp.
layout
Data movement
Pad, reshape, transpose, concatenate, gather, resize, strided slice, tile, broadcast, scatter, mirror pad, and sequence/slice-update utilities.
Classifier tail APIs plus dtype conversion and requantization utilities.
sequence
Sequence
LSTM, GRU, and SVDF functions for temporal workloads.
Convolution Functions
Section titled “Convolution Functions”Conv2D, depthwise, transpose convolution, wrappers, and buffer helpers. 126 functions.
| Function | Summary | Module |
|---|---|---|
arm_convolve_1_x_n_f16 |
1xN convolution, dispatch by layout. | Convolution Functions |
arm_convolve_1_x_n_f16_acc16 |
1xN convolution, dispatch by layout. | Convolution Functions |
arm_convolve_1_x_n_f16_get_buffer_size |
Get the buffer size required by 1xN convolution. | Convolution Functions |
arm_convolve_1_x_n_f32 |
1xN convolution, dispatch by layout. | Convolution Functions |
arm_convolve_1_x_n_f32_get_buffer_size |
Get the buffer size required by 1xN convolution. | Convolution Functions |
arm_convolve_1_x_n_nhwc_f16 |
1xN convolution, NHWC layout. | Convolution Functions |
arm_convolve_1_x_n_nhwc_f16_acc16 |
1xN convolution, NHWC layout. | Convolution Functions |
arm_convolve_1_x_n_nhwc_f32 |
1xN convolution, NHWC layout. | Convolution Functions |
arm_convolve_1_x_n_s4 |
1xn convolution for s4 weights | arm_nnfunctions.h |
arm_convolve_1_x_n_s4_get_buffer_size |
Get the required additional buffer size for 1xn convolution. | arm_nnfunctions.h |
arm_convolve_1_x_n_s8 |
1xn convolution | arm_nnfunctions.h |
arm_convolve_1_x_n_s8_get_buffer_size |
Get the required additional buffer size for 1xn convolution. | arm_nnfunctions.h |
arm_convolve_1x1_f16 |
1x1 convolution, dispatch by layout. | Convolution Functions |
arm_convolve_1x1_f16_acc16 |
1x1 convolution, dispatch by layout. | Convolution Functions |
arm_convolve_1x1_f16_get_buffer_size |
Get the buffer size required by 1x1 convolution. | Convolution Functions |
arm_convolve_1x1_f32 |
1x1 convolution, dispatch by layout. | Convolution Functions |
arm_convolve_1x1_f32_get_buffer_size |
Get the buffer size required by 1x1 convolution. | Convolution Functions |
arm_convolve_1x1_nhwc_f16 |
1x1 convolution, NHWC layout. | Convolution Functions |
arm_convolve_1x1_nhwc_f16_acc16 |
1x1 convolution, NHWC layout. | Convolution Functions |
arm_convolve_1x1_nhwc_f32 |
1x1 convolution, NHWC layout. | Convolution Functions |
arm_convolve_1x1_out_s8 |
Optimised convolution for 1x1 output images (shape of BX1x1xCOUT) for 8x8 computations. | arm_nnfunctions.h |
arm_convolve_1x1_out_s8_get_buffer_size |
Get the required scratch buffer size for armconvolve1x1outs8(). | arm_nnfunctions.h |
arm_convolve_1x1_s16_ns_np_nd |
Pointwise s16 convolution function: no stride, no padding, no dilation. | arm_nnfunctions.h |
arm_convolve_1x1_s4 |
s4 version for 1x1 convolution with support for non-unity stride values | arm_nnfunctions.h |
arm_convolve_1x1_s4_fast |
Fast s4 version for 1x1 convolution (non-square shape). | arm_nnfunctions.h |
arm_convolve_1x1_s4_fast_get_buffer_size |
Get the required buffer size for armconvolve1x1s4fast. | arm_nnfunctions.h |
arm_convolve_1x1_s8 |
s8 version for 1x1 convolution with support for non-unity stride values | arm_nnfunctions.h |
arm_convolve_1x1_s8_fast |
Fast s8 version for 1x1 convolution (non-square shape). | arm_nnfunctions.h |
arm_convolve_1x1_s8_fast_get_buffer_size |
Get the required buffer size for armconvolve1x1s8fast. | arm_nnfunctions.h |
arm_convolve_even_s4 |
Basic s4 convolution function with a requirement of even number of kernels. | arm_nnfunctions.h |
arm_convolve_even_s4_get_buffer_size |
Get the required buffer size for armconvolveevens4. | arm_nnfunctions.h |
arm_convolve_f16 |
Convolution, dispatch by layout. | Convolution Functions |
arm_convolve_f16_acc16 |
Convolution, dispatch by layout. | Convolution Functions |
arm_convolve_f16_get_buffer_size |
Get the temporary buffer size required by convolution. | Convolution Functions |
arm_convolve_f32 |
Convolution, dispatch by layout. | Convolution Functions |
arm_convolve_f32_get_buffer_size |
Get the temporary buffer size required by convolution. | Convolution Functions |
arm_convolve_nhwc_f16 |
Convolution, NHWC layout. | Convolution Functions |
arm_convolve_nhwc_f16_acc16 |
Convolution, NHWC layout. | Convolution Functions |
arm_convolve_nhwc_f32 |
Convolution, NHWC layout. | Convolution Functions |
arm_convolve_s16 |
Basic s16 convolution function. | arm_nnfunctions.h |
arm_convolve_s16_fast_small_kernel |
armconvolves16fastsmallkernel function. | arm_nnfunctions.h |
arm_convolve_s16_get_buffer_size |
Get the required buffer size for s16 convolution function. | arm_nnfunctions.h |
arm_convolve_s16_group_ch_mult_1 |
s16 grouped convolution optimized for the case where filterdims->c == 1 and inputch == outputch (channel multiplier = 1). | arm_nnfunctions.h |
arm_convolve_s4 |
Basic s4 convolution function. | arm_nnfunctions.h |
arm_convolve_s4_get_buffer_size |
Get the required buffer size for s4 convolution function. | arm_nnfunctions.h |
arm_convolve_s8 |
Basic s8 convolution function. | arm_nnfunctions.h |
arm_convolve_s8_3x3_c16_s1 |
s8 3x3 convolution over 16 input channels with unit stride. | arm_nnfunctions.h |
arm_convolve_s8_get_buffer_size |
Get the required buffer size for s8 convolution function. | arm_nnfunctions.h |
arm_convolve_s8_get_buffer_size_mve |
Get the required buffer size for armconvolves8 for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_convolve_s8_get_weights_sum_size |
Get the required buffer size for s8 convolution and depthwise convolution weight sum. | arm_nnfunctions.h |
arm_convolve_s8_small_cin |
s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model. | arm_nnfunctions.h |
arm_convolve_weight_sum |
Pre-computes per-output-channel weight sums for a standard convolution. | arm_nnfunctions.h |
arm_convolve_wrapper_f16 |
Convolution wrapper using the CMSIS-NN baseline path. | Convolution Functions |
arm_convolve_wrapper_f16_acc16 |
Convolution wrapper using the CMSIS-NN baseline path. | Convolution Functions |
arm_convolve_wrapper_f16_get_buffer_size |
Get the buffer size required by the convolution wrapper. | Convolution Functions |
arm_convolve_wrapper_f32 |
Convolution wrapper using the CMSIS-NN baseline path. | Convolution Functions |
arm_convolve_wrapper_f32_get_buffer_size |
Get the buffer size required by the convolution wrapper. | Convolution Functions |
arm_convolve_wrapper_s16 |
s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution. | arm_nnfunctions.h |
arm_convolve_wrapper_s16_get_buffer_size |
Get the required buffer size for armconvolvewrappers16. | arm_nnfunctions.h |
arm_convolve_wrapper_s16_get_buffer_size_dsp |
Get the required buffer size for armconvolvewrappers16 for for processors with DSP extension. | arm_nnfunctions.h |
arm_convolve_wrapper_s16_get_buffer_size_mve |
Get the required buffer size for armconvolvewrappers16 for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_convolve_wrapper_s4 |
s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution. | arm_nnfunctions.h |
arm_convolve_wrapper_s4_get_buffer_size |
Get the required buffer size for armconvolvewrappers4. | arm_nnfunctions.h |
arm_convolve_wrapper_s4_get_buffer_size_dsp |
Get the required buffer size for armconvolvewrappers4 for processors with DSP extension. | arm_nnfunctions.h |
arm_convolve_wrapper_s4_get_buffer_size_mve |
Get the required buffer size for armconvolvewrappers4 for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_convolve_wrapper_s8 |
s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution. | arm_nnfunctions.h |
arm_convolve_wrapper_s8_get_buffer_size |
Get the required buffer size for armconvolvewrappers8. | arm_nnfunctions.h |
arm_convolve_wrapper_s8_get_buffer_size_dsp |
Get the required buffer size for armconvolvewrappers8 for processors with DSP extension. | arm_nnfunctions.h |
arm_convolve_wrapper_s8_get_buffer_size_mve |
Get the required buffer size for armconvolvewrappers8 for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_depthwise_conv_3x3_s8 |
Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below). | arm_nnfunctions.h |
arm_depthwise_conv_f16 |
Depthwise convolution, dispatch by layout. | Convolution Functions |
arm_depthwise_conv_f16_acc16 |
Depthwise convolution, dispatch by layout. | Convolution Functions |
arm_depthwise_conv_f16_get_buffer_size |
Get the temporary buffer size required by depthwise convolution. | Convolution Functions |
arm_depthwise_conv_f32 |
Depthwise convolution, dispatch by layout. | Convolution Functions |
arm_depthwise_conv_f32_get_buffer_size |
Get the temporary buffer size required by depthwise convolution. | Convolution Functions |
arm_depthwise_conv_fast_s16 |
Optimized s16 depthwise convolution function with constraint that inchannel equals outchannel. | arm_nnfunctions.h |
arm_depthwise_conv_fast_s16_get_buffer_size |
Get the required buffer size for optimized s16 depthwise convolution function with constraint that inchannel equals outchannel. | arm_nnfunctions.h |
arm_depthwise_conv_s16 |
Basic s16 depthwise convolution function that doesn’t have any constraints on the input dimensions. | arm_nnfunctions.h |
arm_depthwise_conv_s4 |
Basic s4 depthwise convolution function that doesn’t have any constraints on the input dimensions. | arm_nnfunctions.h |
arm_depthwise_conv_s4_opt |
Optimized s4 depthwise convolution function with constraint that inchannel equals outchannel. | arm_nnfunctions.h |
arm_depthwise_conv_s4_opt_get_buffer_size |
Get the required buffer size for optimized s4 depthwise convolution function with constraint that inchannel equals outchannel. | arm_nnfunctions.h |
arm_depthwise_conv_s8 |
Basic s8 depthwise convolution function that doesn’t have any constraints on the input dimensions. | arm_nnfunctions.h |
arm_depthwise_conv_s8_opt |
Optimized s8 depthwise convolution function with constraint that inchannel equals outchannel. | arm_nnfunctions.h |
arm_depthwise_conv_s8_opt_3x3 |
s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of armdepthwiseconvs8opt(), for layers in its gate. | arm_nnfunctions.h |
arm_depthwise_conv_s8_opt_3x3_c64_s1 |
armdepthwiseconvs8opt3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps. | arm_nnfunctions.h |
arm_depthwise_conv_s8_opt_3x3_get_buffer_size |
Get the scratch size in bytes of armdepthwiseconvs8opt3x3() and armdepthwiseconvs8opt3x3c64s1(). | arm_nnfunctions.h |
arm_depthwise_conv_s8_opt_channelwise |
The channel-vectorized path of armdepthwiseconvs8opt() on its own, without the planar attempt. | arm_nnfunctions.h |
arm_depthwise_conv_s8_opt_get_buffer_size |
Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel. | arm_nnfunctions.h |
arm_depthwise_conv_s8_opt_planar |
The planar path of armdepthwiseconvs8opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane. | arm_nnfunctions.h |
arm_depthwise_conv_s8_opt_planar_supported |
Whether armdepthwiseconvs8opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels. | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_f16 |
Depthwise convolution wrapper using the CMSIS-NN baseline path. | Convolution Functions |
arm_depthwise_conv_wrapper_f16_acc16 |
Depthwise convolution wrapper using the CMSIS-NN baseline path. | Convolution Functions |
arm_depthwise_conv_wrapper_f16_get_buffer_size |
Get the buffer size required by the depthwise convolution wrapper. | Convolution Functions |
arm_depthwise_conv_wrapper_f32 |
Depthwise convolution wrapper using the CMSIS-NN baseline path. | Convolution Functions |
arm_depthwise_conv_wrapper_f32_get_buffer_size |
Get the buffer size required by the depthwise convolution wrapper. | Convolution Functions |
arm_depthwise_conv_wrapper_s16 |
Wrapper function to pick the right optimized s16 depthwise convolution function. | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s16_get_buffer_size |
Get size of additional buffer required by armdepthwiseconvwrappers16(). | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp |
Get size of additional buffer required by armdepthwiseconvwrappers16() for processors with DSP extension. | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s16_get_buffer_size_mve |
Get size of additional buffer required by armdepthwiseconvwrappers16() for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s4 |
Wrapper function to pick the right optimized s4 depthwise convolution function. | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s4_get_buffer_size |
Get size of additional buffer required by armdepthwiseconvwrappers4(). | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp |
Get size of additional buffer required by armdepthwiseconvwrappers4() for processors with DSP extension. | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s4_get_buffer_size_mve |
Get size of additional buffer required by armdepthwiseconvwrappers4() for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s8 |
Wrapper function to pick the right optimized s8 depthwise convolution function. | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s8_get_buffer_size |
Get size of additional buffer required by armdepthwiseconvwrappers8(). | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp |
Get size of additional buffer required by armdepthwiseconvwrappers8() for processors with DSP extension. | arm_nnfunctions.h |
arm_depthwise_conv_wrapper_s8_get_buffer_size_mve |
Get size of additional buffer required by armdepthwiseconvwrappers8() for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_depthwise_convolve_weight_sum |
Pre-computes per-channel weight sums for a depthwise convolution. | arm_nnfunctions.h |
arm_depthwise_nhwc_conv_f16 |
Depthwise convolution, NHWC layout. | Convolution Functions |
arm_depthwise_nhwc_conv_f16_acc16 |
Depthwise convolution, NHWC layout. | Convolution Functions |
arm_depthwise_nhwc_conv_f32 |
Depthwise convolution, NHWC layout. | Convolution Functions |
arm_transpose_conv_f16 |
Transpose convolution, dispatch by layout. | Convolution Functions |
arm_transpose_conv_f16_get_buffer_size |
Get the temporary buffer size required by transpose convolution. | Convolution Functions |
arm_transpose_conv_f16_get_reverse_conv_buffer_size |
Get the reverse-convolution workspace size used by transpose convolution helpers. | Convolution Functions |
arm_transpose_conv_f32 |
Transpose convolution, dispatch by layout. | Convolution Functions |
arm_transpose_conv_f32_get_buffer_size |
Get the temporary buffer size required by transpose convolution. | Convolution Functions |
arm_transpose_conv_f32_get_reverse_conv_buffer_size |
Get the reverse-convolution workspace size used by transpose convolution helpers. | Convolution Functions |
arm_transpose_conv_nhwc_f16 |
Transpose convolution, NHWC layout. | Convolution Functions |
arm_transpose_conv_nhwc_f32 |
Transpose convolution, NHWC layout. | Convolution Functions |
arm_transpose_conv_s8 |
Basic s8 transpose convolution function. | arm_nnfunctions.h |
arm_transpose_conv_s8_get_buffer_size |
Get the required buffer size for ctx in s8 transpose conv function. | arm_nnfunctions.h |
arm_transpose_conv_s8_get_buffer_size_mve |
Get size of additional buffer required by armtransposeconvs8() for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_transpose_conv_s8_get_reverse_conv_buffer_size |
Get the required buffer size for outputctx in s8 transpose conv function. | arm_nnfunctions.h |
arm_transpose_conv_wrapper_f16 |
Transpose convolution wrapper using the CMSIS-NN baseline path. | Convolution Functions |
arm_transpose_conv_wrapper_f32 |
Transpose convolution wrapper using the CMSIS-NN baseline path. | Convolution Functions |
arm_transpose_conv_wrapper_s8 |
Wrapper to select optimal transposed convolution algorithm depending on parameters. | arm_nnfunctions.h |
Fully-Connected Layer Functions
Section titled “Fully-Connected Layer Functions”Dense layers, batch matmul paths, and scratch sizing helpers. 33 functions.
| Function | Summary | Module |
|---|---|---|
arm_batch_matmul_f16 |
Batched matrix multiplication. | Fully-connected Layer Functions |
arm_batch_matmul_f16_get_buffer_size |
Get the temporary buffer size required by batched matrix multiplication. | Fully-connected Layer Functions |
arm_batch_matmul_f32 |
Batched matrix multiplication. | Fully-connected Layer Functions |
arm_batch_matmul_f32_get_buffer_size |
Get the temporary buffer size required by batched matrix multiplication. | Fully-connected Layer Functions |
arm_batch_matmul_s16 |
Batch matmul function with 16 bit input and output. | arm_nnfunctions.h |
arm_batch_matmul_s8 |
Batch matmul function with 8 bit input and output. | arm_nnfunctions.h |
arm_batch_matmul_s8_get_buffer_size |
Get size of the scratch buffer required by armbatchmatmuls8(). | arm_nnfunctions.h |
arm_batch_matmul_s8_get_buffer_size_dsp |
Get size of the scratch buffer required by armbatchmatmuls8() for processors with DSP extension. | arm_nnfunctions.h |
arm_batch_matmul_s8_get_buffer_size_mve |
Get size of the scratch buffer required by armbatchmatmuls8() for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_fully_connected_f16 |
Fully connected layer, dispatch by layout. | Fully-connected Layer Functions |
arm_fully_connected_f16_acc16 |
Fully connected layer, dispatch by layout. | Fully-connected Layer Functions |
arm_fully_connected_f16_get_buffer_size |
Get the temporary buffer size required by the fully connected layer. | Fully-connected Layer Functions |
arm_fully_connected_f32 |
Fully connected layer, dispatch by layout. | Fully-connected Layer Functions |
arm_fully_connected_f32_get_buffer_size |
Get the temporary buffer size required by the fully connected layer. | Fully-connected Layer Functions |
arm_fully_connected_nhwc_f16 |
Fully connected layer, NHWC layout. | Fully-connected Layer Functions |
arm_fully_connected_nhwc_f16_acc16 |
Fully connected layer, NHWC layout. | Fully-connected Layer Functions |
arm_fully_connected_nhwc_f32 |
Fully connected layer, NHWC layout. | Fully-connected Layer Functions |
arm_fully_connected_per_channel_s16 |
Basic s16 Fully Connected function using per channel quantization. | arm_nnfunctions.h |
arm_fully_connected_per_channel_s16_get_buffer_size |
Get size of additional buffer required by armfullyconnectedperchannels16(). | arm_nnfunctions.h |
arm_fully_connected_per_channel_s16_get_buffer_size_dsp |
Get size of additional buffer required by armfullyconnectedperchannels16() for processors with DSP extension. | arm_nnfunctions.h |
arm_fully_connected_per_channel_s16_get_buffer_size_mve |
Get size of additional buffer required by armfullyconnectedperchannels16() for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_fully_connected_per_channel_s8 |
Basic s8 Fully Connected function using per channel quantization. | arm_nnfunctions.h |
arm_fully_connected_s16 |
Basic s16 Fully Connected function. | arm_nnfunctions.h |
arm_fully_connected_s16_get_buffer_size |
Get size of additional buffer required by armfullyconnecteds16(). | arm_nnfunctions.h |
arm_fully_connected_s16_get_buffer_size_dsp |
Get size of additional buffer required by armfullyconnecteds16() for processors with DSP extension. | arm_nnfunctions.h |
arm_fully_connected_s16_get_buffer_size_mve |
Get size of additional buffer required by armfullyconnecteds16() for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_fully_connected_s4 |
Basic s4 Fully Connected function. | arm_nnfunctions.h |
arm_fully_connected_s8 |
Basic s8 Fully Connected function. | arm_nnfunctions.h |
arm_fully_connected_s8_get_buffer_size |
Get size of additional buffer required by armfullyconnecteds8(). | arm_nnfunctions.h |
arm_fully_connected_s8_get_buffer_size_dsp |
Get size of additional buffer required by armfullyconnecteds8() for processors with DSP extension. | arm_nnfunctions.h |
arm_fully_connected_s8_get_buffer_size_mve |
Get size of additional buffer required by armfullyconnecteds8() for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_fully_connected_wrapper_s16 |
s16 Fully Connected layer wrapper function | arm_nnfunctions.h |
arm_fully_connected_wrapper_s8 |
s8 Fully Connected layer wrapper function | arm_nnfunctions.h |
Elementwise Functions
Section titled “Elementwise Functions”Add, sub, mul, square difference, min/max, batch norm, select, and arithmetic glue. 65 functions.
| Function | Summary | Module |
|---|---|---|
arm_abs_s16 |
s16 elementwise absolute value | arm_nnfunctions.h |
arm_abs_s8 |
s8 elementwise absolute value | arm_nnfunctions.h |
arm_add_s16 |
s16 elementwise add of two tensors with support for broadcasting. | arm_nnfunctions.h |
arm_add_s8 |
s8 elementwise add of two tensors with support for broadcasting. | arm_nnfunctions.h |
arm_add_scalar_s16 |
s16 elementwise add of scalar and vector | arm_nnfunctions.h |
arm_add_scalar_s8 |
s8 elementwise add of scalar and vector | arm_nnfunctions.h |
arm_batch_norm_f16 |
Apply batch normalization. | NNSupport |
arm_batch_norm_f32 |
Apply batch normalization. | NNSupport |
arm_elementwise_add_broadcast_f16 |
Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp. | Elementwise Functions |
arm_elementwise_add_broadcast_f32 |
Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp. | Elementwise Functions |
arm_elementwise_add_f16 |
Elementwise add with optional output clamp. | Elementwise Functions |
arm_elementwise_add_f32 |
Elementwise add with optional output clamp. | Elementwise Functions |
arm_elementwise_add_fp16 |
Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate armelementwiseaddf16(). | Elementwise Functions |
arm_elementwise_add_s16 |
s16 elementwise add of two vectors | arm_nnfunctions.h |
arm_elementwise_add_s8 |
s8 elementwise add of two vectors | arm_nnfunctions.h |
arm_elementwise_mul_broadcast_f16 |
Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp. | Elementwise Functions |
arm_elementwise_mul_broadcast_f32 |
Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp. | Elementwise Functions |
arm_elementwise_mul_f16 |
Elementwise multiply with optional output clamp. | Elementwise Functions |
arm_elementwise_mul_f32 |
Elementwise multiply with optional output clamp. | Elementwise Functions |
arm_elementwise_mul_s16 |
s16 elementwise multiplication | arm_nnfunctions.h |
arm_elementwise_mul_s8 |
s8 elementwise multiplication | arm_nnfunctions.h |
arm_elementwise_prelu_s16 |
Elementwise S16 PReLU activation function. | arm_nnfunctions.h |
arm_elementwise_prelu_s8 |
Elementwise S8 PReLU activation function. | arm_nnfunctions.h |
arm_elementwise_squared_difference_f16 |
Elementwise squared difference of two float16 vectors. | Elementwise Functions |
arm_elementwise_squared_difference_s16 |
s16 elementwise squared difference of two vectors. | arm_nnfunctions.h |
arm_elementwise_squared_difference_s8 |
s8 elementwise squared difference of two vectors. | arm_nnfunctions.h |
arm_elementwise_sub_broadcast_f16 |
Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp. | Elementwise Functions |
arm_elementwise_sub_broadcast_f32 |
Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp. | Elementwise Functions |
arm_elementwise_sub_f16 |
Elementwise subtract with optional output clamp. | Elementwise Functions |
arm_elementwise_sub_f32 |
Elementwise subtract with optional output clamp. | Elementwise Functions |
arm_elementwise_sub_s16 |
s16 elementwise subtract of two vectors | arm_nnfunctions.h |
arm_elementwise_sub_s8 |
s8 elementwise subtract of two vectors | arm_nnfunctions.h |
arm_maximum_f16 |
Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. | Elementwise Functions |
arm_maximum_f32 |
Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. | Elementwise Functions |
arm_maximum_s16 |
s16 elementwise maximum w/ support for broadcasting and scalar inputs. | arm_nnfunctions.h |
arm_maximum_s8 |
s8 elementwise maximum w/ support for broadcasting and scalar inputs. | arm_nnfunctions.h |
arm_minimum_f16 |
Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. | Elementwise Functions |
arm_minimum_f32 |
Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape. | Elementwise Functions |
arm_minimum_s16 |
s16 elementwise minimum w/ support for broadcasting and scalar inputs. | arm_nnfunctions.h |
arm_minimum_s8 |
s8 elementwise minimum w/ support for broadcasting and scalar inputs. | arm_nnfunctions.h |
arm_mul_s16 |
s16 elementwise multiplication of two tensors with support for broadcasting. | arm_nnfunctions.h |
arm_mul_s8 |
s8 elementwise multiplication of two tensors with support for broadcasting. | arm_nnfunctions.h |
arm_mul_scalar_s16 |
s16 elementwise multiplication of scalar and vector | arm_nnfunctions.h |
arm_mul_scalar_s8 |
s8 elementwise multiplication of scalar and vector | arm_nnfunctions.h |
arm_nn_abs_f16 |
Elementwise absolute value. | Elementwise Functions |
arm_nn_abs_f32 |
Elementwise absolute value. | Elementwise Functions |
arm_nn_sqrt_f16 |
Elementwise square root of a float16 tensor. | Elementwise Functions |
arm_nn_sqrt_f32 |
Elementwise square root. | Elementwise Functions |
arm_rsqrt_f16 |
Elementwise reciprocal square root of a float16 tensor, 1 / sqrt(x). | Elementwise Functions |
arm_rsqrt_f32 |
Elementwise reciprocal square root, 1 / sqrt(x). | Elementwise Functions |
arm_rsqrt_s16_per_op |
INT16 reciprocal square root using a per-operator LUT. | arm_nnfunctions.h |
arm_rsqrt_s16_universal |
INT16 reciprocal square root using a shared universal LUT. | arm_nnfunctions.h |
arm_select_v2_s16 |
SELECTV2 with broadcast for int16 tensors. | arm_nnfunctions.h |
arm_select_v2_s8 |
SELECTV2 with broadcast for int8 tensors. | arm_nnfunctions.h |
arm_sqrt_s16 |
s16 elementwise square root using piecewise LUT with linear interpolation | arm_nnfunctions.h |
arm_sqrt_s16_tablefree |
s16 elementwise square root without a lookup table | arm_nnfunctions.h |
arm_sqrt_s8 |
s8 elementwise square root | arm_nnfunctions.h |
arm_squared_difference_s16 |
s16 elementwise squared difference of two tensors with support for broadcasting. | arm_nnfunctions.h |
arm_squared_difference_s8 |
s8 elementwise squared difference of two tensors with support for broadcasting. | arm_nnfunctions.h |
arm_squared_difference_scalar_s16 |
s16 elementwise squared difference of scalar and vector. | arm_nnfunctions.h |
arm_squared_difference_scalar_s8 |
s8 elementwise squared difference of scalar and vector. | arm_nnfunctions.h |
arm_sub_s16 |
s16 elementwise subtraction of two tensors with support for broadcasting. | arm_nnfunctions.h |
arm_sub_s8 |
s8 elementwise subtraction of two tensors with support for broadcasting. | arm_nnfunctions.h |
arm_sub_scalar_s16 |
s16 elementwise subtract of scalar and vector (scalar - vector) | arm_nnfunctions.h |
arm_sub_scalar_s8 |
s8 elementwise subtract of scalar and vector (scalar - vector) | arm_nnfunctions.h |
Basic Math and Reduction
Section titled “Basic Math and Reduction”Argmin/argmax, min/max reductions, comparisons, means, vector sums, and where. 40 functions.
| Function | Summary | Module |
|---|---|---|
arm_argmax_f16 |
Returns the first maximum’s axis-relative INT32 index for a f16 tensor. | Reduction Functions |
arm_argmax_f32 |
Returns the first maximum’s axis-relative INT32 index for a f32 tensor. | Reduction Functions |
arm_argmax_s16 |
Compute ArgMax indices of an s16 tensor along a specific axis. | arm_nnfunctions.h |
arm_argmax_s8 |
Compute ArgMax indices of an s8 tensor along a specific axis. | arm_nnfunctions.h |
arm_argmin_f16 |
Returns the first minimum’s axis-relative INT32 index for a f16 tensor. | Reduction Functions |
arm_argmin_f32 |
Returns the first minimum’s axis-relative INT32 index for a f32 tensor. | Reduction Functions |
arm_argmin_s16 |
Compute ArgMin indices of an s16 tensor along a specific axis. | arm_nnfunctions.h |
arm_argmin_s8 |
Compute ArgMin indices of an s8 tensor along a specific axis. | arm_nnfunctions.h |
arm_comparison_s16 |
s16 elementwise comparison with support for broadcasting. | arm_nnfunctions.h |
arm_comparison_s8 |
s8 elementwise comparison with support for broadcasting. | arm_nnfunctions.h |
arm_equal_s16 |
s16 elementwise equality comparison with support for broadcasting. | arm_nnfunctions.h |
arm_equal_s8 |
s8 elementwise equality comparison with support for broadcasting. | arm_nnfunctions.h |
arm_greater_equal_s16 |
s16 elementwise greater-or-equal comparison with support for broadcasting. | arm_nnfunctions.h |
arm_greater_equal_s8 |
s8 elementwise greater-or-equal comparison with support for broadcasting. | arm_nnfunctions.h |
arm_greater_s16 |
s16 elementwise greater-than comparison with support for broadcasting. | arm_nnfunctions.h |
arm_greater_s8 |
s8 elementwise greater-than comparison with support for broadcasting. | arm_nnfunctions.h |
arm_less_equal_s16 |
s16 elementwise less-or-equal comparison with support for broadcasting. | arm_nnfunctions.h |
arm_less_equal_s8 |
s8 elementwise less-or-equal comparison with support for broadcasting. | arm_nnfunctions.h |
arm_less_s16 |
s16 elementwise less-than comparison with support for broadcasting. | arm_nnfunctions.h |
arm_less_s8 |
s8 elementwise less-than comparison with support for broadcasting. | arm_nnfunctions.h |
arm_mean_s16 |
Computes the mean of the input tensor along the specified axis. | arm_nnfunctions.h |
arm_mean_s8 |
Computes the mean of the input tensor along the specified axis. | arm_nnfunctions.h |
arm_nn_mean_f16 |
Computes the mean of a float16 tensor along the specified axes. | Reduction Functions |
arm_nn_mean_f32 |
Computes the mean of a float32 tensor along the specified axes. | Reduction Functions |
arm_not_equal_s16 |
s16 elementwise inequality comparison with support for broadcasting. | arm_nnfunctions.h |
arm_not_equal_s8 |
s8 elementwise inequality comparison with support for broadcasting. | arm_nnfunctions.h |
arm_reduce_max_f16 |
Reduces a f16 NHWC tensor to its maximum along a binary axis mask. | Reduction Functions |
arm_reduce_max_f32 |
Reduces a f32 NHWC tensor to its maximum along a binary axis mask. | Reduction Functions |
arm_reduce_max_s16 |
Computes the max of the input tensor along the specified axis. | arm_nnfunctions.h |
arm_reduce_max_s8 |
Computes the max of the input tensor along the specified axis. | arm_nnfunctions.h |
arm_reduce_min_f16 |
Reduces a f16 NHWC tensor to its minimum along a binary axis mask. | Reduction Functions |
arm_reduce_min_f32 |
Reduces a f32 NHWC tensor to its minimum along a binary axis mask. | Reduction Functions |
arm_reduce_min_s16 |
Computes the min of the input tensor along the specified axis. | arm_nnfunctions.h |
arm_reduce_min_s8 |
Computes the min of the input tensor along the specified axis. | arm_nnfunctions.h |
arm_reduce_sum_f16 |
Computes the sum of the input tensor along the specified axes. | Reduction Functions |
arm_reduce_sum_f32 |
Computes the sum of the input tensor along the specified axes. | Reduction Functions |
arm_vector_sum_s8 |
Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s32 biasdata. | arm_nnfunctions.h |
arm_vector_sum_s8_s64 |
Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s64 biasdata. | arm_nnfunctions.h |
arm_where_s16 |
WHERE operator: return coordinates of non-zero elements in condition (int16). | arm_nnfunctions.h |
arm_where_s8 |
WHERE operator: return coordinates of non-zero elements in condition. | arm_nnfunctions.h |
Activation Functions
Section titled “Activation Functions”ReLU, LeakyReLU, PReLU, Hard-Swish, Logistic, Tanh, and clamp. 27 functions.
| Function | Summary | Module |
|---|---|---|
arm_clamp_s16 |
S16 clamp function. | arm_nnfunctions.h |
arm_clamp_s8 |
S8 clamp function. | arm_nnfunctions.h |
arm_hard_swish_compat_s8 |
S8 Hard-Swish activation function (compatibility version). | arm_nnfunctions.h |
arm_hard_swish_f16 |
Hard swish activation for float16 data. | Activation Functions |
arm_hard_swish_f32 |
Hard swish activation for float32 data. | Activation Functions |
arm_hard_swish_precise_s16 |
S16 Hard-Swish activation function (precise version). | arm_nnfunctions.h |
arm_hard_swish_precise_s8 |
S8 Hard-Swish activation function (precise version). | arm_nnfunctions.h |
arm_leaky_relu_s16 |
S16 Leaky ReLU activation function. | arm_nnfunctions.h |
arm_leaky_relu_s8 |
S8 Leaky ReLU activation function. | arm_nnfunctions.h |
arm_logistic_s16 |
Logistic activation function for s16. | arm_nnfunctions.h |
arm_nn_activation_f16 |
Elementwise activation. | Activation Functions |
arm_nn_activation_f32 |
Elementwise activation. | Activation Functions |
arm_nn_activation_s16 |
s16 neural network activation function using direct table look-up | arm_nnfunctions.h |
arm_prelu_f16 |
Parametric ReLU for float32 data. | Activation Functions |
arm_prelu_f32 |
Parametric ReLU for float32 data. | Activation Functions |
arm_prelu_s16 |
S16 PReLU activation function. | arm_nnfunctions.h |
arm_prelu_s8 |
S8 PReLU activation function. | arm_nnfunctions.h |
arm_prelu_scalar_s16 |
Scalar S16 PReLU activation function. | arm_nnfunctions.h |
arm_prelu_scalar_s8 |
Scalar S8 PReLU activation function. | arm_nnfunctions.h |
arm_relu_generic_s16 |
S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc. | arm_nnfunctions.h |
arm_relu_generic_s8 |
S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc. | arm_nnfunctions.h |
arm_relu_q15 |
Q15 RELU function. | arm_nnfunctions.h |
arm_relu_q7 |
Q7 RELU function. | arm_nnfunctions.h |
arm_relu_s16 |
S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767. | arm_nnfunctions.h |
arm_relu_s8 |
S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127. | arm_nnfunctions.h |
arm_relu6_q7 |
Q7 RELU6 function. | arm_nnfunctions.h |
arm_tanh_s16 |
Tanh activation function for s16. | arm_nnfunctions.h |
Data Movement
Section titled “Data Movement”Pad, reshape, transpose, concatenate, gather, resize, strided slice, tile, broadcast, scatter, mirror pad, and sequence/slice-update utilities. 77 functions.
| Function | Summary | Module |
|---|---|---|
arm_batch_to_space_nd_s16 |
Batch to Space ND function for s16 data type. | arm_nnfunctions.h |
arm_batch_to_space_nd_s8 |
Batch to Space ND function for s8 data type. | arm_nnfunctions.h |
arm_broadcast_to_s16 |
Broadcast an int16 tensor to a target shape. | arm_nnfunctions.h |
arm_broadcast_to_s8 |
Broadcast an int8 tensor to a target shape. | arm_nnfunctions.h |
arm_concatenation_f16 |
Concatenate float32 tensors of any rank along one axis. | NNSupport |
arm_concatenation_f16_w |
Concatenate tensors along the W axis. | NNSupport |
arm_concatenation_f16_x |
Concatenate tensors along the X axis. | NNSupport |
arm_concatenation_f16_y |
Concatenate tensors along the Y axis. | NNSupport |
arm_concatenation_f16_z |
Concatenate tensors along the Z axis. | NNSupport |
arm_concatenation_f32 |
Concatenate float32 tensors of any rank along one axis. | NNSupport |
arm_concatenation_f32_w |
Concatenate tensors along the W axis. | NNSupport |
arm_concatenation_f32_x |
Concatenate tensors along the X axis. | NNSupport |
arm_concatenation_f32_y |
Concatenate tensors along the Y axis. | NNSupport |
arm_concatenation_f32_z |
Concatenate tensors along the Z axis. | NNSupport |
arm_concatenation_s16 |
int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis | arm_nnfunctions.h |
arm_concatenation_s32 |
int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis | arm_nnfunctions.h |
arm_concatenation_s8 |
int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis | arm_nnfunctions.h |
arm_concatenation_s8_w |
int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to… | arm_nnfunctions.h |
arm_concatenation_s8_x |
int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate. | arm_nnfunctions.h |
arm_concatenation_s8_y |
int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate. | arm_nnfunctions.h |
arm_concatenation_s8_z |
int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate. | arm_nnfunctions.h |
arm_depth_to_space_s16 |
Depth to Space function for s16 data type. | arm_nnfunctions.h |
arm_depth_to_space_s8 |
Depth to Space function for s8 data type. | arm_nnfunctions.h |
arm_dynamic_update_slice_s16 |
Update a slice of an int16 operand tensor at runtime-determined indices. | arm_nnfunctions.h |
arm_dynamic_update_slice_s8 |
Update a slice of an int8 operand tensor at runtime-determined indices. | arm_nnfunctions.h |
arm_gather_f16 |
Gather contiguous slices along an axis. | Gather Functions: |
arm_gather_f32 |
Gather contiguous slices along an axis. | Gather Functions: |
arm_gather_nd_f16 |
Gather contiguous slices using coordinate tuples. | Gather Functions: |
arm_gather_nd_f32 |
Gather contiguous slices using coordinate tuples. | Gather Functions: |
arm_gather_nd_s16 |
Gathernd slices for int16 tensors. | arm_nnfunctions.h |
arm_gather_nd_s8 |
Gathernd slices for int8 tensors. | arm_nnfunctions.h |
arm_gather_s16 |
Gather elements along an axis for int16 tensors. | arm_nnfunctions.h |
arm_gather_s8 |
Gather elements along an axis for int8 tensors. | arm_nnfunctions.h |
arm_mirror_pad_s16 |
Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode). | arm_nnfunctions.h |
arm_mirror_pad_s8 |
Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode). | arm_nnfunctions.h |
arm_nn_fill_f16 |
Fill a float16 vector with one value; bit copy of value, NaN payload included. | Elementwise Functions |
arm_nn_fill_f32 |
Fill a float32 vector with one value. | Elementwise Functions |
arm_pack_f16 |
Stack float32 tensors of equal shape along a new axis (TFLite PACK). | NNSupport |
arm_pack_f32 |
Stack float32 tensors of equal shape along a new axis (TFLite PACK). | NNSupport |
arm_pad_f16 |
Pad a tensor with a constant value. | Pad Layer Functions: |
arm_pad_f32 |
Pad a tensor with a constant value. | Pad Layer Functions: |
arm_pad_s16 |
Expands the size of the input by adding constant values before and after the data, in all dimensions. | arm_nnfunctions.h |
arm_pad_s8 |
Expands the size of the input by adding constant values before and after the data, in all dimensions. | arm_nnfunctions.h |
arm_reshape_f16 |
Reshape by copying data without changing element order. | NNSupport |
arm_reshape_f32 |
Reshape by copying data without changing element order. | NNSupport |
arm_reshape_s8 |
Reshape a s8 vector into another with different shape. | arm_nnfunctions.h |
arm_resize_nearest_neighbor_f16 |
Nearest-neighbor resize of a float32 NHWC tensor. | Reshape Functions |
arm_resize_nearest_neighbor_f16_get_buffer_size |
Scratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16(). | Reshape Functions |
arm_resize_nearest_neighbor_f32 |
Nearest-neighbor resize of a float32 NHWC tensor. | Reshape Functions |
arm_resize_nearest_neighbor_f32_get_buffer_size |
Scratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16(). | Reshape Functions |
arm_resize_nearest_neighbor_s16 |
Nearest neighbor resize function for s16 data. | arm_nnfunctions.h |
arm_resize_nearest_neighbor_s8 |
Nearest neighbor resize function for s8 data. | arm_nnfunctions.h |
arm_reverse_sequence_s16 |
Reverse variable-length sequences along a dimension for int16. | arm_nnfunctions.h |
arm_reverse_sequence_s8 |
Reverse variable-length sequences along a dimension for int8. | arm_nnfunctions.h |
arm_scatter_nd_s16 |
Scatter updates into a zero-initialized output tensor for int16. | arm_nnfunctions.h |
arm_scatter_nd_s8 |
Scatter updates into a zero-initialized output tensor for int8. | arm_nnfunctions.h |
arm_space_to_batch_nd_s16 |
Space to Batch ND function for s16 data type. | arm_nnfunctions.h |
arm_space_to_batch_nd_s8 |
Space to Batch ND function for s8 data type. | arm_nnfunctions.h |
arm_space_to_depth_s16 |
Space to Depth function for s16 data type. | arm_nnfunctions.h |
arm_space_to_depth_s8 |
Space to Depth function for s8 data type. | arm_nnfunctions.h |
arm_split_f16 |
Split a float32 tensor of any rank into several tensors along one axis. | Elementwise Functions |
arm_split_f32 |
Split a float32 tensor of any rank into several tensors along one axis. | NNSupport |
arm_split_s16 |
int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis | arm_nnfunctions.h |
arm_split_s8 |
int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis | arm_nnfunctions.h |
arm_strided_slice_f16 |
Strided slice for float32 data (pure copy, TensorFlow Lite compatible). | Elementwise Functions |
arm_strided_slice_f32 |
Strided slice for float32 data (pure copy, TensorFlow Lite compatible). | Slicing Functions: |
arm_strided_slice_s16 |
Strided slice function for int16 data. | arm_nnfunctions.h |
arm_strided_slice_s32 |
Strided slice function for int32 data. | arm_nnfunctions.h |
arm_strided_slice_s8 |
Strided slice function for int8 data. | arm_nnfunctions.h |
arm_tile_s16 |
Tile an int16 tensor along each dimension. | arm_nnfunctions.h |
arm_tile_s8 |
Tile an int8 tensor along each dimension. | arm_nnfunctions.h |
arm_transpose_f16 |
Transpose a floating-point tensor. | NNSupport |
arm_transpose_f32 |
Transpose a floating-point tensor. | NNSupport |
arm_transpose_s16 |
Basic s16 transpose function. | arm_nnfunctions.h |
arm_transpose_s8 |
Basic transpose function. | arm_nnfunctions.h |
arm_unpack_f16 |
Unstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK). | NNSupport |
arm_unpack_f32 |
Unstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK). | NNSupport |
Classifier Tail
Section titled “Classifier Tail”Classifier tail APIs plus dtype conversion and requantization utilities. 27 functions.
| Function | Summary | Module |
|---|---|---|
arm_avg_pool_f16 |
Average pooling. | Pooling Functions |
arm_avg_pool_f32 |
Average pooling. | Pooling Functions |
arm_avgpool_s16 |
s16 average pooling function. | arm_nnfunctions.h |
arm_avgpool_s16_get_buffer_size |
Get the required buffer size for S16 average pooling function. | arm_nnfunctions.h |
arm_avgpool_s16_get_buffer_size_dsp |
Get the required buffer size for S16 average pooling function for processors with DSP extension. | arm_nnfunctions.h |
arm_avgpool_s16_get_buffer_size_mve |
Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_avgpool_s8 |
s8 average pooling function. | arm_nnfunctions.h |
arm_avgpool_s8_get_buffer_size |
Get the required buffer size for S8 average pooling function. | arm_nnfunctions.h |
arm_avgpool_s8_get_buffer_size_dsp |
Get the required buffer size for S8 average pooling function for processors with DSP extension. | arm_nnfunctions.h |
arm_avgpool_s8_get_buffer_size_mve |
Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_dequantize_f16_f32 |
Widen a float16 vector to float32. | Quantization Functions: |
arm_dequantize_s16_f32 |
Dequantize an int16t array back to floating-point format. | arm_nnfunctions.h |
arm_dequantize_s8_f32 |
Dequantize an int8t array back to floating-point format. | arm_nnfunctions.h |
arm_max_pool_f16 |
Max pooling. | Pooling Functions |
arm_max_pool_f32 |
Max pooling. | Pooling Functions |
arm_max_pool_s16 |
s16 max pooling function. | arm_nnfunctions.h |
arm_max_pool_s8 |
s8 max pooling function. | arm_nnfunctions.h |
arm_quantize_f32_s16 |
Quantize a floating-point array into int16t format. | arm_nnfunctions.h |
arm_quantize_f32_s8 |
Quantize a floating-point array into int8t format. | arm_nnfunctions.h |
arm_requantize_s16_s16 |
Requantize an int16t array to another int16t range with a different scale. | arm_nnfunctions.h |
arm_requantize_s8_s8 |
Requantize an int8t array to another int8t range with a different scale. | arm_nnfunctions.h |
arm_softmax_f16 |
Softmax using the float-native API signature. | Softmax Functions |
arm_softmax_f32 |
Softmax using the float-native API signature. | Softmax Functions |
arm_softmax_s16 |
S16 softmax function. | arm_nnfunctions.h |
arm_softmax_s8 |
S8 softmax function. | arm_nnfunctions.h |
arm_softmax_s8_s16 |
S8 to s16 softmax function. | arm_nnfunctions.h |
arm_softmax_u8 |
U8 softmax function. | arm_nnfunctions.h |
Sequence Functions
Section titled “Sequence Functions”LSTM, GRU, and SVDF functions for temporal workloads. 31 functions.
| Function | Summary | Module |
|---|---|---|
arm_gru_unidirectional_f16 |
Unidirectional GRU layer for float16 input, output and state. | LSTM Layer Functions |
arm_gru_unidirectional_f16_temp1_get_buffer_size |
Get size of the temp1 scratch buffer required by armgruunidirectionalf16(). | LSTM Layer Functions |
arm_gru_unidirectional_f32 |
Unidirectional GRU layer for float32 input, output and state. | LSTM Layer Functions |
arm_gru_unidirectional_f32_temp1_get_buffer_size |
Get size of the temp1 scratch buffer required by armgruunidirectionalf32(). | LSTM Layer Functions |
arm_lstm_unidirectional_f16 |
Unidirectional LSTM inference. | LSTM Layer Functions |
arm_lstm_unidirectional_f16_temp1_get_buffer_size |
Get size of the temp1 scratch buffer required by armlstmunidirectionalf16(). | LSTM Layer Functions |
arm_lstm_unidirectional_f16_temp2_get_buffer_size |
Get size of the temp2 scratch buffer required by armlstmunidirectionalf16(). | LSTM Layer Functions |
arm_lstm_unidirectional_f32 |
Unidirectional LSTM inference. | LSTM Layer Functions |
arm_lstm_unidirectional_f32_temp1_get_buffer_size |
Get size of the temp1 scratch buffer required by armlstmunidirectionalf32(). | LSTM Layer Functions |
arm_lstm_unidirectional_f32_temp2_get_buffer_size |
Get size of the temp2 scratch buffer required by armlstmunidirectionalf32(). | LSTM Layer Functions |
arm_lstm_unidirectional_s16 |
LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias. | arm_nnfunctions.h |
arm_lstm_unidirectional_s16_temp1_get_buffer_size |
Get size of the temp1 scratch buffer required by armlstmunidirectionals16(). | arm_nnfunctions.h |
arm_lstm_unidirectional_s16_temp2_get_buffer_size |
Get size of the temp2 scratch buffer required by armlstmunidirectionals16(). | arm_nnfunctions.h |
arm_lstm_unidirectional_s8 |
LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias. | arm_nnfunctions.h |
arm_lstm_unidirectional_s8_temp1_get_buffer_size |
Get size of the temp1 scratch buffer required by armlstmunidirectionals8(). | arm_nnfunctions.h |
arm_lstm_unidirectional_s8_temp2_get_buffer_size |
Get size of the temp2 scratch buffer required by armlstmunidirectionals8(). | arm_nnfunctions.h |
arm_svdf_f16 |
Stateful singular value decomposition filter, float16 variant. | SVDF Functions |
arm_svdf_f16_input_ctx_get_buffer_size |
Get size of the inputctx staging buffer required by armsvdff16(). | SVDF Functions |
arm_svdf_f16_output_ctx_get_buffer_size |
Get size of the outputctx staging buffer required by armsvdff16(). | SVDF Functions |
arm_svdf_f32 |
Stateful singular value decomposition filter. | SVDF Functions |
arm_svdf_f32_input_ctx_get_buffer_size |
Get size of the inputctx staging buffer required by armsvdff32(). | SVDF Functions |
arm_svdf_f32_output_ctx_get_buffer_size |
Get size of the outputctx staging buffer required by armsvdff32(). | SVDF Functions |
arm_svdf_s8 |
s8 SVDF function with 8 bit state tensor and 8 bit time weights | arm_nnfunctions.h |
arm_svdf_s8_get_buffer_size |
Get size of the kernel-sum buffer required by armsvdfs8(). | arm_nnfunctions.h |
arm_svdf_s8_get_buffer_size_dsp |
Get size of the kernel-sum buffer required by armsvdfs8() for processors with DSP extension. | arm_nnfunctions.h |
arm_svdf_s8_get_buffer_size_mve |
Get size of the kernel-sum buffer required by armsvdfs8() for Arm(R) Helium Architecture case. | arm_nnfunctions.h |
arm_svdf_s8_input_ctx_get_buffer_size |
Get size of the inputctx staging buffer required by armsvdfs8(). | arm_nnfunctions.h |
arm_svdf_s8_output_ctx_get_buffer_size |
Get size of the outputctx staging buffer required by armsvdfs8(). | arm_nnfunctions.h |
arm_svdf_state_s16_s8 |
s8 SVDF function with 16 bit state tensor and 16 bit time weights | arm_nnfunctions.h |
arm_svdf_state_s16_s8_input_ctx_get_buffer_size |
Get size of the inputctx staging buffer required by armsvdfstates16s8(). | arm_nnfunctions.h |
arm_svdf_state_s16_s8_output_ctx_get_buffer_size |
Get size of the outputctx staging buffer required by armsvdfstates16s8(). | arm_nnfunctions.h |
Types and support
Section titled “Types and support”Structs, enums, macros and typedefs stay on the module pages rather than being repeated here. Open heliaCORE for the full tree.