Operator catalog
Explore 70 registered operators. Search by name or restriction, combine family and data-type filters, then expand an operator for its requirements.
Data types describe individual tensors, not complete A8W8 or A16W8 inference paths. Support also depends on tensor roles, shapes, attributes and target; conversion validates your actual graph.
70 of 70 operators
Accelerator
ETHOS_UAcceleratorint8
Restrictions
- Requires a target platform that declares an NPU.
- The command-stream tensor must be a constant carrying its data.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
Activation
HARD_SWISHActivationint8int16float16float32
Restrictions
- Input and output must share one dtype and one shape.
- Float HARD_SWISH requires positive dimensions and at most INT32_MAX elements.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.32.0
- Context buffer
- no_ctx
LEAKY_RELUActivationint8int16
Restrictions
- Input and output must both use a supported dtype.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
LOGISTICActivationint8int16float16float32
Restrictions
- Quantized input and output must carry quantization parameters.
- An int8 output must use zero-point -128.
- int16 input and output tensors must use zero-point 0.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
PRELUActivationint8float16float32
Restrictions
- Input, alpha and output tensors must share one dtype.
- The float alpha shape must broadcast to the input: every alpha dimension equals the input dimension or is 1.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.31.0
- Context buffer
- no_ctx
RELUActivationint8int16float16float32
Restrictions
- Float RELU supports the RELU and RELU6 activation types only.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
SOFTMAXActivationint8int16float16float32
Restrictions
- An int8 input produces an int8 or int16 output; an int16 input produces an int16 output.
- Float softmax requires matching input and output dtypes.
- Float softmax requires beta == 1.0: the CMSIS-NN float kernels take no beta.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
TANHActivationint8int16float16float32
Restrictions
- Quantized input and output must carry quantization parameters.
- int16 input and output tensors must use zero-point 0.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
Comparison
COMPARISONComparisonint8int16
Restrictions
- Both inputs must share one dtype and carry a quantization scale and zero-point.
- The output tensor must be boolean.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- ignored
SELECT_V2Comparisonint8int16
Restrictions
- The condition tensor must be boolean; the x, y and output tensors must be int8 or int16 sharing one dtype.
- Rank must be 8 or lower and the output shape must be the broadcast of the input shapes.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
WHEREComparisonint8int16
Restrictions
- The condition tensor must be int8 or int16 with rank 1 to 8.
- The output tensor must be int64 and must match the shape the kernel derives from the condition.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
Convolution
CONV_2DConvolutionint8int16float16float32
Restrictions
- Input and output must share one dtype, and float weights and bias must use that same dtype.
- Weight scales must be per-tensor or one per output channel on the output-channel axis (0 for OHWI).
- Upscale factors must be 1 or greater, and upscaling is supported on the int8 kernel path only.
- The weights and bias tensors must be constants so the weight sums can be precomputed.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- conditional
DEPTHWISE_CONV_2DConvolutionint8int16float16float32
Restrictions
- Input and output must share one dtype, and float weights and bias must use that same dtype.
- Weight scales must be per-tensor or one per output channel on the output-channel axis (3 for 1HWO).
- The weights and bias tensors must be constants, with rank 3 or 4, so the weight sums can be precomputed.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- conditional
TRANSPOSE_CONVConvolutionint8float16float32
Restrictions
- Input and output must share one dtype.
- The int8 path requires int8 weights and, when present, an int32 bias.
- The float path requires weights and any bias to use the input dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- conditional
Dense
FULLY_CONNECTEDDenseint8int16float16float32
Restrictions
- Input and output must share one supported dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- conditional
Elementwise
ABSElementwiseint8int16float16float32
Restrictions
- Input and output must have the same dtype and the same shape.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.31.0
- Context buffer
- no_ctx
ADDElementwiseint8int16float16float32
Restrictions
- Both inputs and the output must share one dtype.
- Float ADD is elementwise only: rank-4-or-lower shapes may differ only by leading size-one dimensions; shapes above rank 4 must be identical.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
MAXIMUMElementwiseint8int16float16float32
Restrictions
- Takes exactly two input tensors and one output tensor, all of a supported dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- ignored
MINIMUMElementwiseint8int16float16float32
Restrictions
- Takes exactly two input tensors and one output tensor, all of a supported dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- ignored
MULElementwiseint8int16float16float32
Restrictions
- Both inputs and the output must share one dtype.
- Float MUL broadcasts rank-4-or-lower tensors after left-padding them to NHWC; equal shapes of any rank retain the flat kernel.
- Float MUL requires positive dimensions and at most INT32_MAX elements.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.32.0
- Context buffer
- no_ctx
RSQRTElementwiseint8int16float16float32
Restrictions
- Input and output must share one dtype and one shape.
- Quantized tensors must carry a scale and a zero-point, and int16 tensors must use zero-point 0.
- Float RSQRT requires positive dimensions and at most INT32_MAX elements.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.33.0
- Context buffer
- no_ctx
SQRTElementwiseint8int16float16float32
Restrictions
- Input and output must share one dtype and one shape.
- Quantized tensors must carry a scale and a zero-point.
- float16 SQRT requires positive dimensions and at most INT32_MAX elements.
- Float dependency
- float16
- Minimum ns-cmsis-nn for float
- 7.33.0
- Context buffer
- no_ctx
SQUARED_DIFFERENCEElementwiseint8int16
Restrictions
- Takes exactly two input tensors and one output tensor, all in the same integer dtype.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
SUBElementwiseint8int16float16float32
Restrictions
- Both inputs and the output must share one dtype.
- Float SUB supports elementwise shapes, a scalar operand, or NumPy/TFLite broadcasting of operands of rank 4 or lower, compared in the left-padded NHWC view.
- Scalar-broadcast float SUB supports the NONE fused activation only.
- Broadcast dimensions must be positive and the element counts must fit int32.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.32.0
- Context buffer
- no_ctx
Matmul
BATCH_MATMULMatmulint8int16float16float32
Restrictions
- Both operands and the output must share one dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- conditional
Pooling
AVERAGE_POOL_2DPoolingint8int16float16float32
Restrictions
- Takes exactly one input tensor and one output tensor.
- Input and output must both use a supported dtype; pooling does not convert element types.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- conditional
MAX_POOL_2DPoolingint8int16float16float32
Restrictions
- Input and output must both use a supported dtype; pooling does not convert element types.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- ignored
Quantization
DEQUANTIZEQuantizationint8int16float16float32
Restrictions
- Accepts int8, int16 or float16 input and produces float32 output.
- Input and output shapes must match.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
QUANTIZEQuantizationint8int16float32
Restrictions
- Accepts float32 input to int8/int16 output, or int8-to-int8 and int16-to-int16 requantization.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
Reduction
ARG_MAXReductionint8int16float16float32
Restrictions
- The output tensor must be int32.
- The axis attribute must be in range for the input rank.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
ARG_MINReductionint8int16float16float32
Restrictions
- The output tensor must be int32.
- The axis attribute must be in range for the input rank.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
MEANReductionint8int16float16float32
Restrictions
- Input and output must share one supported dtype.
- Input rank must be 1 to 4 with positive dimensions.
- Reduction axes must be constant and in range; integer MEAN requires at least one axis.
- Float input element count must fit int32.
- Float output shape must match the reduction with reduced dimensions kept or removed.
- Integer output shape must match keep_dims; when unspecified, kept or removed dimensions are accepted.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.32.0
- Context buffer
- no_ctx
REDUCE_MAXReductionint8int16float16float32
Restrictions
- Input and output must share one dtype; there is no requantization.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.34.0
- Context buffer
- no_ctx
REDUCE_MINReductionint8int16float16float32
Restrictions
- Input and output must share one dtype; there is no requantization.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.34.0
- Context buffer
- no_ctx
SUMReductionint8int16float16float32
Restrictions
- Input and output must share one dtype; integer tensors must carry quantization parameters.
- The axes tensor must be constant int32 with axes in the input rank's range.
- Input and output rank must be at most 4.
- The output shape must match the reduction with reduced dimensions kept or removed.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.31.0
- Context buffer
- no_ctx
Stateful
ASSIGN_VARIABLEStatefulint8int16float16float32
Restrictions
- Resource and assigned tensors must have the same numeric dtype.
- Float16 tensors require a platform with FP16 kernel support.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
GRUStatefulfloat16
Restrictions
- Takes two inputs (sequence and initial state) and two outputs (sequence and final state).
- Every tensor must be float16, and the weight tensors must be constants.
- Sequence tensors must be rank 3 with positive dimensions, batch-major, and batch size 1.
- Only reset_after=True is supported.
- The initial state and the final state must be distinct tensors.
- Float dependency
- float16
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
READ_VARIABLEStatefulint8int16float16float32
Restrictions
- Resource and output tensors must have the same numeric dtype.
- Float16 tensors require a platform with FP16 kernel support.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
SVDFStatefulint8float16float32
Restrictions
- Kernel rank limits are 1..32767 for integer paths and 1..2147483647 for float paths.
- Input and output must share one dtype.
- On the float path the bias and weight tensors must use the input dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- required
UNIDIRECTIONAL_SEQUENCE_LSTMStatefulint8float16float32
Restrictions
- Only the standard Keras LSTM tensor subset is supported, with TANH cell activation.
- Projection clipping, diagonal recurrent tensors and asymmetric_quantize_inputs are not supported.
- Input and output must share one dtype.
- Weight and bias tensors must be constants whose shapes follow the input and hidden sizes.
- The output state and the cell state must each hold batch size times hidden size elements.
- On the int8 path the weights must be symmetric (zero-point 0) and the output state must be int8.
- On the int8 path the output state must carry the quantization of the primary output.
- On the int8 path the cell state must be int16 with zero-point 0 and a power-of-two scale.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
Tensor
BATCH_TO_SPACE_NDTensorint8int16
Restrictions
- Input rank must be 3 or 4, with one block_shape entry per spatial dimension.
- Block shape entries must be non-zero and the batch dimension must divide by each of them.
- Crops must be non-negative and every computed output dimension must be positive.
- Input and output must share one dtype; there is no requantization.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
BROADCAST_TOTensorint8int16
Restrictions
- Input and output tensors must be int8 or int16 with rank 1 to 8.
- The shape tensor must be a constant rank-1 int32 or int64 tensor holding only positive dimensions.
- The output shape must equal the shape tensor and must be a valid broadcast of the input shape.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
CONCATENATIONTensorint8int16int32float16float32
Restrictions
- Takes between one and ten input tensors, all with the same non-zero rank as the output.
- Float concatenation requires 4-D NHWC tensors, and every input must use the output dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
DEPTH_TO_SPACETensorint8int16
Restrictions
- Input and output must share one dtype; there is no requantization.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
DILATETensorint8int16float16float32
Restrictions
- Requires DILATE options and a constant int32 dilations tensor with one entry per input dimension.
- Input rank must be 1 to 4 and every dilation factor must be 1 or greater.
- Each output dimension must equal (input dimension - 1) * dilation + 1.
- The padding value must be a finite constant scalar in the input dtype.
- Input and output must share one dtype.
- When both tensors carry quantization metadata, first scales must be close (numpy.isclose) and first zero-points must be equal.
- Float dependency
- float16
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
DYNAMIC_UPDATE_SLICETensorint8int16
Restrictions
- Operand, update and output tensors must be int8 or int16 and share one dtype.
- Operand rank must be 1 to 8, the update rank must match it, and no update dimension may exceed the operand.
- start_indices must be a rank-1 int32 or int64 tensor with one entry per operand dimension.
- The output shape must equal the operand shape.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
EXPAND_DIMSTensorint8int16int32float16float32
Restrictions
- Takes exactly one input tensor and one output tensor.
- Moves tensor bytes without numeric conversion; aliased input/output buffers require no copy.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
FILLTensorint8int16float16float32
Restrictions
- The fill value must use the output dtype: the emitted cast reinterprets it rather than requantizing it.
- Float dependency
- float16
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
GATHERTensorint8int16float16float32
Restrictions
- The indices tensor must be int32 and the output dtype must match the data tensor.
- The axis must be in range for the input rank.
- The output rank must equal input rank plus indices rank minus batch_dims minus one.
- Float tensors must have positive dimensions; data rank is 1 to 4 and indices/output rank is 0 to 4.
- Each float-path tensor buffer must fit within INT32_MAX bytes.
- Float batch_dims must be nonnegative, at most axis and indices rank, and less than data rank.
- Float data and indices must have matching leading batch dimensions.
- The float output shape must match the shape inferred from data, indices, axis and batch_dims.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.34.0
- Context buffer
- no_ctx
GATHER_NDTensorint8int16float16float32
Restrictions
- The indices tensor must be int32 and the output dtype must match the params tensor.
- Float option shapes must match the actual data, indices and output tensor shapes.
- Float tensors must have positive dimensions; data/indices rank is 1 to 4 and output rank is 0 to 4.
- Each float-path tensor buffer must fit within INT32_MAX bytes.
- Float index depth must not exceed data rank.
- The float output shape must equal indices.shape[:-1] followed by data.shape[index_depth:].
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.34.0
- Context buffer
- no_ctx
MIRROR_PADTensorint8int16
Restrictions
- Input and output tensors must be int8 or int16, rank 1 to 8, sharing one dtype, scale and zero-point.
- int16 tensors must use zero-point 0.
- The paddings tensor must have shape [rank, 2] and hold non-negative values.
- Paddings must stay within the input-size limits of the selected mirror mode.
- The output shape must equal the padded input shape.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
PACKTensorint8int16int32float16float32
Restrictions
- Takes at least one input, and the input count must equal both values_count and the size of the pack axis.
- Every input must use the output dtype and must equal the output shape with the pack axis removed.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.33.1
- Context buffer
- no_ctx
PADTensorint8int16float16float32
Restrictions
- Pre- and post-padding values must be given.
- Input and output must share one supported dtype.
- Quantized input and output must share one scale.
- For quantized tensors, the constant padding value must fit the output dtype range.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
RESHAPETensorint8int16float16float32
Restrictions
- Input and output must share one dtype and the same total byte count.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
RESIZE_BILINEARTensorint8int16
Restrictions
- Input and output tensors must be int8 or int16 with rank 4 or lower, sharing one scale and zero-point.
- The size tensor must be a constant int32 tensor of exactly two elements matching the output height and width.
- half_pixel_centers and align_corners must not both be true.
- The output batch and channel dimensions must match the input.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
RESIZE_NEAREST_NEIGHBORTensorint8int16float16float32
Restrictions
- The size tensor must be a constant int32 tensor of two elements: the kernel precomputes its coordinate map.
- The size tensor values must match the output height and width.
- Rank must be 1 to 4 with positive dimensions, and the output batch and channels must match the input.
- The element count and the coordinate map must both fit int32 indexing.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.33.0
- Context buffer
- required
REVERSE_SEQUENCETensorint8int16
Restrictions
- Input and output tensors must be int8 or int16, sharing one dtype and one shape, with rank 1 to 8.
- seq_lengths must be a rank-1 int32 tensor as long as the batch dimension.
- seq_lengths values must be non-negative and no larger than the sequence dimension.
- seq_dim and batch_dim must be non-negative, different, and in range for the input rank.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
REVERSE_V2Tensorint8int16float16float32
Restrictions
- Input and output must share one dtype and one shape, with rank 1 or higher.
- Axes must be unique once normalized and in range for the input rank.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
SCATTER_NDTensorint8int16
Restrictions
- Updates and output tensors must be int8 or int16 and share one dtype.
- The indices tensor and the shape tensor must be int32, and the shape tensor must be a constant.
- The output shape must equal the contents of the shape tensor.
- The index depth must be between 1 and the output rank.
- Indices and updates must share their leading dimensions, and updates must match the output suffix dimensions.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
SHAPETensorint8int16float16float32
Restrictions
No restrictions declared in the catalog. Conversion still validates the graph.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
SLICETensorint8int16float16float32
Restrictions
- Input rank must be 4 or lower.
- The output must use the input dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
SPACE_TO_BATCH_NDTensorint8int16
Restrictions
- Input rank must be 3 or 4, with one non-zero block_shape entry per spatial dimension.
- Each padded spatial dimension must divide by its block_shape entry.
- Input and output must share one dtype; there is no requantization.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
SPACE_TO_DEPTHTensorint8int16
Restrictions
- Input and output must share one dtype; there is no requantization.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
SPLITTensorint8int16int32float16float32
Restrictions
- Takes one input tensor and at least one output, and split_lengths must have one entry per output.
- The axis must be in range, every split length must be positive, and the lengths must sum to the axis size.
- Every output must use the input dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.33.1
- Context buffer
- no_ctx
SQUEEZETensorint8int16float16float32
Restrictions
- SQUEEZE only drops size-1 dimensions: input and output must share one dtype and the same byte count.
- Float dependency
- float16
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
STRIDED_SLICETensorint8int16int32float16float32
Restrictions
- Input rank must be 4 or lower.
- The output must use the input dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
TILETensorint8int16
Restrictions
- Input and output tensors must be int8 or int16 and share one dtype.
- The multipliers tensor must be int32, or a constant int64 tensor whose values fit int32.
- There must be one non-negative multiplier per input dimension.
- The output shape must equal the input shape multiplied elementwise by the multipliers.
- Float dependency
- None declared
- Minimum ns-cmsis-nn for float
- Not applicable
- Context buffer
- no_ctx
TRANSPOSETensorint8int16float16float32
Restrictions
- Input and output must both use a supported dtype.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
UNPACKTensorint8int16int32float16float32
Restrictions
- Takes one input tensor and at least one output, and the output count must equal both num and the axis size.
- Every output must use the input dtype and must equal the input shape with the unpack axis removed.
- Float dependency
- float16, float32
- Minimum ns-cmsis-nn for float
- 7.33.1
- Context buffer
- no_ctx
ZEROS_LIKETensorint8int16int32float16float32
Restrictions
- Input and output must have the same shape and the same dtype.
- Float dependency
- float16
- Minimum ns-cmsis-nn for float
- 7.35.0
- Context buffer
- no_ctx
Reading a row
Section titled “Reading a row”| Column | What it says |
|---|---|
| Data types | A summary of data tensor types, including conversion input/output types. This is not every legal input/output pairing or the type of every index, mask, weight or state tensor. Byte-copy and metadata rows may admit additional types. |
| Float dependency | The types for which the lowering declares a gated float API or C type dependency. This may cover arithmetic or only data movement. no means no such dependency is declared; it does not rule out integer-library helpers or inline float code. Neither value proves SIMD optimization. |
| Minimum ns-cmsis-nn | The oldest library the float path builds against. The module-wide floor is 7.35.0; an operator with no float support says —. |
| Context buffer | How the lowering uses cmsis_nn_context::buf: no_ctx, ignored, required, conditional. |
| Restrictions | Documented shape, quantization, attribute, platform and kernel limits. Individual validators enforce the applicable subset. |
For example, DEQUANTIZE takes int8, int16 or float16 and produces float32; QUANTIZE takes float32 to int8/int16 or requantizes within the same integer width. ARG_MIN/MAX produce int32 indices, comparisons produce bool, and WHERE produces int64 indices. The int8 LSTM path has int8 output state and int16 cell state; an int16 state is not an int16 inference path.
Lowering can call heliaCORE, emit inline arithmetic or lookup loops, copy bytes, reuse an aliased buffer, or invoke the ETHOS-U accelerator. EXPAND_DIMS copies bytes or becomes a no-op; DILATE and RESIZE_BILINEAR emit inline loops. A heliaCORE call may itself select a vector, DSP or scalar path. Check the generated source and the linked library configuration when comparing optimized coverage.
Each context buffer size must fit 0..2147483647 bytes. Dimensions and sizing parameters passed through signed 32-bit kernel fields must also fit those fields. Whether zero dimensions are permitted depends on the operator, converter and selected kernel or sizer; do not infer one universal lower bound. Products can overflow even when individual dimensions fit.
Full reference tables by family
Accelerator operators
Section titled “Accelerator operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
ETHOS_U |
int8 |
no | — | no_ctx |
Requires a target platform that declares an NPU. The command-stream tensor must be a constant carrying its data. |
Activation operators
Section titled “Activation operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
HARD_SWISH |
int8, int16, float16, float32 |
float16, float32 |
7.32.0 |
no_ctx |
Input and output must share one dtype and one shape. Float HARD_SWISH requires positive dimensions and at most INT32_MAX elements. |
LEAKY_RELU |
int8, int16 |
no | — | no_ctx |
Input and output must both use a supported dtype. |
LOGISTIC |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Quantized input and output must carry quantization parameters. An int8 output must use zero-point -128. int16 input and output tensors must use zero-point 0. |
PRELU |
int8, float16, float32 |
float16, float32 |
7.31.0 |
no_ctx |
Input, alpha and output tensors must share one dtype. The float alpha shape must broadcast to the input: every alpha dimension equals the input dimension or is 1. |
RELU |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Float RELU supports the RELU and RELU6 activation types only. |
SOFTMAX |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
An int8 input produces an int8 or int16 output; an int16 input produces an int16 output. Float softmax requires matching input and output dtypes. Float softmax requires beta == 1.0: the CMSIS-NN float kernels take no beta. |
TANH |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Quantized input and output must carry quantization parameters. int16 input and output tensors must use zero-point 0. |
Comparison operators
Section titled “Comparison operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
COMPARISON |
int8, int16 |
no | — | ignored |
Both inputs must share one dtype and carry a quantization scale and zero-point. The output tensor must be boolean. |
SELECT_V2 |
int8, int16 |
no | — | no_ctx |
The condition tensor must be boolean; the x, y and output tensors must be int8 or int16 sharing one dtype. Rank must be 8 or lower and the output shape must be the broadcast of the input shapes. |
WHERE |
int8, int16 |
no | — | no_ctx |
The condition tensor must be int8 or int16 with rank 1 to 8. The output tensor must be int64 and must match the shape the kernel derives from the condition. |
Convolution operators
Section titled “Convolution operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
CONV_2D |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
conditional |
Input and output must share one dtype, and float weights and bias must use that same dtype. Weight scales must be per-tensor or one per output channel on the output-channel axis (0 for OHWI). Upscale factors must be 1 or greater, and upscaling is supported on the int8 kernel path only. The weights and bias tensors must be constants so the weight sums can be precomputed. |
DEPTHWISE_CONV_2D |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
conditional |
Input and output must share one dtype, and float weights and bias must use that same dtype. Weight scales must be per-tensor or one per output channel on the output-channel axis (3 for 1HWO). The weights and bias tensors must be constants, with rank 3 or 4, so the weight sums can be precomputed. |
TRANSPOSE_CONV |
int8, float16, float32 |
float16, float32 |
7.35.0 |
conditional |
Input and output must share one dtype. The int8 path requires int8 weights and, when present, an int32 bias. The float path requires weights and any bias to use the input dtype. |
Dense operators
Section titled “Dense operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
FULLY_CONNECTED |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
conditional |
Input and output must share one supported dtype. |
Elementwise operators
Section titled “Elementwise operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
ABS |
int8, int16, float16, float32 |
float16, float32 |
7.31.0 |
no_ctx |
Input and output must have the same dtype and the same shape. |
ADD |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Both inputs and the output must share one dtype. Float ADD is elementwise only: rank-4-or-lower shapes may differ only by leading size-one dimensions; shapes above rank 4 must be identical. |
MAXIMUM |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
ignored |
Takes exactly two input tensors and one output tensor, all of a supported dtype. |
MINIMUM |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
ignored |
Takes exactly two input tensors and one output tensor, all of a supported dtype. |
MUL |
int8, int16, float16, float32 |
float16, float32 |
7.32.0 |
no_ctx |
Both inputs and the output must share one dtype. Float MUL broadcasts rank-4-or-lower tensors after left-padding them to NHWC; equal shapes of any rank retain the flat kernel. Float MUL requires positive dimensions and at most INT32_MAX elements. |
RSQRT |
int8, int16, float16, float32 |
float16, float32 |
7.33.0 |
no_ctx |
Input and output must share one dtype and one shape. Quantized tensors must carry a scale and a zero-point, and int16 tensors must use zero-point 0. Float RSQRT requires positive dimensions and at most INT32_MAX elements. |
SQRT |
int8, int16, float16, float32 |
float16 |
7.33.0 |
no_ctx |
Input and output must share one dtype and one shape. Quantized tensors must carry a scale and a zero-point. float16 SQRT requires positive dimensions and at most INT32_MAX elements. |
SQUARED_DIFFERENCE |
int8, int16 |
no | — | no_ctx |
Takes exactly two input tensors and one output tensor, all in the same integer dtype. |
SUB |
int8, int16, float16, float32 |
float16, float32 |
7.32.0 |
no_ctx |
Both inputs and the output must share one dtype. Float SUB supports elementwise shapes, a scalar operand, or NumPy/TFLite broadcasting of operands of rank 4 or lower, compared in the left-padded NHWC view. Scalar-broadcast float SUB supports the NONE fused activation only. Broadcast dimensions must be positive and the element counts must fit int32. |
Matmul operators
Section titled “Matmul operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
BATCH_MATMUL |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
conditional |
Both operands and the output must share one dtype. |
Pooling operators
Section titled “Pooling operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
AVERAGE_POOL_2D |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
conditional |
Takes exactly one input tensor and one output tensor. Input and output must both use a supported dtype; pooling does not convert element types. |
MAX_POOL_2D |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
ignored |
Input and output must both use a supported dtype; pooling does not convert element types. |
Quantization operators
Section titled “Quantization operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
DEQUANTIZE |
int8, int16, float16, float32 |
no | 7.35.0 |
no_ctx |
Accepts int8, int16 or float16 input and produces float32 output. Input and output shapes must match. |
QUANTIZE |
int8, int16, float32 |
no | 7.35.0 |
no_ctx |
Accepts float32 input to int8/int16 output, or int8-to-int8 and int16-to-int16 requantization. |
Reduction operators
Section titled “Reduction operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
ARG_MAX |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
The output tensor must be int32. The axis attribute must be in range for the input rank. |
ARG_MIN |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
The output tensor must be int32. The axis attribute must be in range for the input rank. |
MEAN |
int8, int16, float16, float32 |
float16, float32 |
7.32.0 |
no_ctx |
Input and output must share one supported dtype. Input rank must be 1 to 4 with positive dimensions. Reduction axes must be constant and in range; integer MEAN requires at least one axis. Float input element count must fit int32. Float output shape must match the reduction with reduced dimensions kept or removed. Integer output shape must match keep_dims; when unspecified, kept or removed dimensions are accepted. |
REDUCE_MAX |
int8, int16, float16, float32 |
float16, float32 |
7.34.0 |
no_ctx |
Input and output must share one dtype; there is no requantization. |
REDUCE_MIN |
int8, int16, float16, float32 |
float16, float32 |
7.34.0 |
no_ctx |
Input and output must share one dtype; there is no requantization. |
SUM |
int8, int16, float16, float32 |
float16, float32 |
7.31.0 |
no_ctx |
Input and output must share one dtype; integer tensors must carry quantization parameters. The axes tensor must be constant int32 with axes in the input rank’s range. Input and output rank must be at most 4. The output shape must match the reduction with reduced dimensions kept or removed. |
Stateful operators
Section titled “Stateful operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
ASSIGN_VARIABLE |
int8, int16, float16, float32 |
no | 7.35.0 |
no_ctx |
Resource and assigned tensors must have the same numeric dtype. Float16 tensors require a platform with FP16 kernel support. |
GRU |
float16 |
float16 |
7.35.0 |
no_ctx |
Takes two inputs (sequence and initial state) and two outputs (sequence and final state). Every tensor must be float16, and the weight tensors must be constants. Sequence tensors must be rank 3 with positive dimensions, batch-major, and batch size 1. Only reset_after=True is supported. The initial state and the final state must be distinct tensors. |
READ_VARIABLE |
int8, int16, float16, float32 |
no | 7.35.0 |
no_ctx |
Resource and output tensors must have the same numeric dtype. Float16 tensors require a platform with FP16 kernel support. |
SVDF |
int8, float16, float32 |
float16, float32 |
7.35.0 |
required |
Kernel rank limits are 1..32767 for integer paths and 1..2147483647 for float paths. Input and output must share one dtype. On the float path the bias and weight tensors must use the input dtype. |
UNIDIRECTIONAL_SEQUENCE_LSTM |
int8, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Only the standard Keras LSTM tensor subset is supported, with TANH cell activation. Projection clipping, diagonal recurrent tensors and asymmetric_quantize_inputs are not supported. Input and output must share one dtype. Weight and bias tensors must be constants whose shapes follow the input and hidden sizes. The output state and the cell state must each hold batch size times hidden size elements. On the int8 path the weights must be symmetric (zero-point 0) and the output state must be int8. On the int8 path the output state must carry the quantization of the primary output. On the int8 path the cell state must be int16 with zero-point 0 and a power-of-two scale. |
Tensor operators
Section titled “Tensor operators”| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
|---|---|---|---|---|---|
BATCH_TO_SPACE_ND |
int8, int16 |
no | — | no_ctx |
Input rank must be 3 or 4, with one block_shape entry per spatial dimension. Block shape entries must be non-zero and the batch dimension must divide by each of them. Crops must be non-negative and every computed output dimension must be positive. Input and output must share one dtype; there is no requantization. |
BROADCAST_TO |
int8, int16 |
no | — | no_ctx |
Input and output tensors must be int8 or int16 with rank 1 to 8. The shape tensor must be a constant rank-1 int32 or int64 tensor holding only positive dimensions. The output shape must equal the shape tensor and must be a valid broadcast of the input shape. |
CONCATENATION |
int8, int16, int32, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Takes between one and ten input tensors, all with the same non-zero rank as the output. Float concatenation requires 4-D NHWC tensors, and every input must use the output dtype. |
DEPTH_TO_SPACE |
int8, int16 |
no | — | no_ctx |
Input and output must share one dtype; there is no requantization. |
DILATE |
int8, int16, float16, float32 |
float16 |
7.35.0 |
no_ctx |
Requires DILATE options and a constant int32 dilations tensor with one entry per input dimension. Input rank must be 1 to 4 and every dilation factor must be 1 or greater. Each output dimension must equal (input dimension - 1) * dilation + 1. The padding value must be a finite constant scalar in the input dtype. Input and output must share one dtype. When both tensors carry quantization metadata, first scales must be close (numpy.isclose) and first zero-points must be equal. |
DYNAMIC_UPDATE_SLICE |
int8, int16 |
no | — | no_ctx |
Operand, update and output tensors must be int8 or int16 and share one dtype. Operand rank must be 1 to 8, the update rank must match it, and no update dimension may exceed the operand. start_indices must be a rank-1 int32 or int64 tensor with one entry per operand dimension. The output shape must equal the operand shape. |
EXPAND_DIMS |
int8, int16, int32, float16, float32 |
no | 7.35.0 |
no_ctx |
Takes exactly one input tensor and one output tensor. Moves tensor bytes without numeric conversion; aliased input/output buffers require no copy. |
FILL |
int8, int16, float16, float32 |
float16 |
7.35.0 |
no_ctx |
The fill value must use the output dtype: the emitted cast reinterprets it rather than requantizing it. |
GATHER |
int8, int16, float16, float32 |
float16, float32 |
7.34.0 |
no_ctx |
The indices tensor must be int32 and the output dtype must match the data tensor. The axis must be in range for the input rank. The output rank must equal input rank plus indices rank minus batch_dims minus one. Float tensors must have positive dimensions; data rank is 1 to 4 and indices/output rank is 0 to 4. Each float-path tensor buffer must fit within INT32_MAX bytes. Float batch_dims must be nonnegative, at most axis and indices rank, and less than data rank. Float data and indices must have matching leading batch dimensions. The float output shape must match the shape inferred from data, indices, axis and batch_dims. |
GATHER_ND |
int8, int16, float16, float32 |
float16, float32 |
7.34.0 |
no_ctx |
The indices tensor must be int32 and the output dtype must match the params tensor. Float option shapes must match the actual data, indices and output tensor shapes. Float tensors must have positive dimensions; data/indices rank is 1 to 4 and output rank is 0 to 4. Each float-path tensor buffer must fit within INT32_MAX bytes. Float index depth must not exceed data rank. The float output shape must equal indices.shape[:-1] followed by data.shape[index_depth:]. |
MIRROR_PAD |
int8, int16 |
no | — | no_ctx |
Input and output tensors must be int8 or int16, rank 1 to 8, sharing one dtype, scale and zero-point. int16 tensors must use zero-point 0. The paddings tensor must have shape [rank, 2] and hold non-negative values. Paddings must stay within the input-size limits of the selected mirror mode. The output shape must equal the padded input shape. |
PACK |
int8, int16, int32, float16, float32 |
float16, float32 |
7.33.1 |
no_ctx |
Takes at least one input, and the input count must equal both values_count and the size of the pack axis. Every input must use the output dtype and must equal the output shape with the pack axis removed. |
PAD |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Pre- and post-padding values must be given. Input and output must share one supported dtype. Quantized input and output must share one scale. For quantized tensors, the constant padding value must fit the output dtype range. |
RESHAPE |
int8, int16, float16, float32 |
no | 7.35.0 |
no_ctx |
Input and output must share one dtype and the same total byte count. |
RESIZE_BILINEAR |
int8, int16 |
no | — | no_ctx |
Input and output tensors must be int8 or int16 with rank 4 or lower, sharing one scale and zero-point. The size tensor must be a constant int32 tensor of exactly two elements matching the output height and width. half_pixel_centers and align_corners must not both be true. The output batch and channel dimensions must match the input. |
RESIZE_NEAREST_NEIGHBOR |
int8, int16, float16, float32 |
float16, float32 |
7.33.0 |
required |
The size tensor must be a constant int32 tensor of two elements: the kernel precomputes its coordinate map. The size tensor values must match the output height and width. Rank must be 1 to 4 with positive dimensions, and the output batch and channels must match the input. The element count and the coordinate map must both fit int32 indexing. |
REVERSE_SEQUENCE |
int8, int16 |
no | — | no_ctx |
Input and output tensors must be int8 or int16, sharing one dtype and one shape, with rank 1 to 8. seq_lengths must be a rank-1 int32 tensor as long as the batch dimension. seq_lengths values must be non-negative and no larger than the sequence dimension. seq_dim and batch_dim must be non-negative, different, and in range for the input rank. |
REVERSE_V2 |
int8, int16, float16, float32 |
no | 7.35.0 |
no_ctx |
Input and output must share one dtype and one shape, with rank 1 or higher. Axes must be unique once normalized and in range for the input rank. |
SCATTER_ND |
int8, int16 |
no | — | no_ctx |
Updates and output tensors must be int8 or int16 and share one dtype. The indices tensor and the shape tensor must be int32, and the shape tensor must be a constant. The output shape must equal the contents of the shape tensor. The index depth must be between 1 and the output rank. Indices and updates must share their leading dimensions, and updates must match the output suffix dimensions. |
SHAPE |
int8, int16, float16, float32 |
no | 7.35.0 |
no_ctx |
None declared. |
SLICE |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Input rank must be 4 or lower. The output must use the input dtype. |
SPACE_TO_BATCH_ND |
int8, int16 |
no | — | no_ctx |
Input rank must be 3 or 4, with one non-zero block_shape entry per spatial dimension. Each padded spatial dimension must divide by its block_shape entry. Input and output must share one dtype; there is no requantization. |
SPACE_TO_DEPTH |
int8, int16 |
no | — | no_ctx |
Input and output must share one dtype; there is no requantization. |
SPLIT |
int8, int16, int32, float16, float32 |
float16, float32 |
7.33.1 |
no_ctx |
Takes one input tensor and at least one output, and split_lengths must have one entry per output. The axis must be in range, every split length must be positive, and the lengths must sum to the axis size. Every output must use the input dtype. |
SQUEEZE |
int8, int16, float16, float32 |
float16 |
7.35.0 |
no_ctx |
SQUEEZE only drops size-1 dimensions: input and output must share one dtype and the same byte count. |
STRIDED_SLICE |
int8, int16, int32, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Input rank must be 4 or lower. The output must use the input dtype. |
TILE |
int8, int16 |
no | — | no_ctx |
Input and output tensors must be int8 or int16 and share one dtype. The multipliers tensor must be int32, or a constant int64 tensor whose values fit int32. There must be one non-negative multiplier per input dimension. The output shape must equal the input shape multiplied elementwise by the multipliers. |
TRANSPOSE |
int8, int16, float16, float32 |
float16, float32 |
7.35.0 |
no_ctx |
Input and output must both use a supported dtype. |
UNPACK |
int8, int16, int32, float16, float32 |
float16, float32 |
7.33.1 |
no_ctx |
Takes one input tensor and at least one output, and the output count must equal both num and the axis size. Every output must use the input dtype and must equal the input shape with the unpack axis removed. |
ZEROS_LIKE |
int8, int16, int32, float16, float32 |
float16 |
7.35.0 |
no_ctx |
Input and output must have the same shape and the same dtype. |