# Operator catalog

Explore 70 registered operators. Search by name or restriction, combine family and data-type filters, then expand an operator for its requirements.

Data types describe individual tensors, not complete A8W8 or A16W8 inference paths. Support also depends on tensor roles, shapes, attributes and target; conversion validates your actual graph.

## Reading a row

| Column | What it says |
| --- | --- |
| Data types | A summary of data tensor types, including conversion input/output types. This is not every legal input/output pairing or the type of every index, mask, weight or state tensor. Byte-copy and metadata rows may admit additional types. |
| Float dependency | The types for which the lowering declares a gated float API or C type dependency. This may cover arithmetic or only data movement. `no` means no such dependency is declared; it does not rule out integer-library helpers or inline float code. Neither value proves SIMD optimization. |
| Minimum ns-cmsis-nn | The oldest library the float path builds against. The module-wide floor is `7.35.0`; an operator with no float support says —. |
| Context buffer | How the lowering uses `cmsis_nn_context::buf`: `no_ctx`, `ignored`, `required`, `conditional`. |
| Restrictions | Documented shape, quantization, attribute, platform and kernel limits. Individual validators enforce the applicable subset. |

For example, DEQUANTIZE takes int8, int16 or float16 and produces float32; QUANTIZE takes float32 to int8/int16 or requantizes within the same integer width. ARG_MIN/MAX produce int32 indices, comparisons produce bool, and WHERE produces int64 indices. The int8 LSTM path has int8 output state and int16 cell state; an int16 state is not an int16 inference path.

Lowering can call heliaCORE, emit inline arithmetic or lookup loops, copy bytes, reuse an aliased buffer, or invoke the ETHOS-U accelerator. EXPAND_DIMS copies bytes or becomes a no-op; DILATE and RESIZE_BILINEAR emit inline loops. A heliaCORE call may itself select a vector, DSP or scalar path. Check the generated source and the linked library configuration when comparing optimized coverage.

Each context buffer size must fit 0..2147483647 bytes. Dimensions and sizing parameters passed through signed 32-bit kernel fields must also fit those fields. Whether zero dimensions are permitted depends on the operator, converter and selected kernel or sizer; do not infer one universal lower bound. Products can overflow even when individual dimensions fit.

Full reference tables by family

## Accelerator operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `ETHOS_U` | `int8` | no | — | `no_ctx` | Requires a target platform that declares an NPU. The command-stream tensor must be a constant carrying its data. |

## Activation operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `HARD_SWISH` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.32.0` | `no_ctx` | Input and output must share one dtype and one shape. Float HARD_SWISH requires positive dimensions and at most INT32_MAX elements. |
| `LEAKY_RELU` | `int8`, `int16` | no | — | `no_ctx` | Input and output must both use a supported dtype. |
| `LOGISTIC` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Quantized input and output must carry quantization parameters. An int8 output must use zero-point -128. int16 input and output tensors must use zero-point 0. |
| `PRELU` | `int8`, `float16`, `float32` | `float16`, `float32` | `7.31.0` | `no_ctx` | Input, alpha and output tensors must share one dtype. The float alpha shape must broadcast to the input: every alpha dimension equals the input dimension or is 1. |
| `RELU` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Float RELU supports the RELU and RELU6 activation types only. |
| `SOFTMAX` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | An int8 input produces an int8 or int16 output; an int16 input produces an int16 output. Float softmax requires matching input and output dtypes. Float softmax requires beta == 1.0: the CMSIS-NN float kernels take no beta. |
| `TANH` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Quantized input and output must carry quantization parameters. int16 input and output tensors must use zero-point 0. |

## Comparison operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `COMPARISON` | `int8`, `int16` | no | — | `ignored` | Both inputs must share one dtype and carry a quantization scale and zero-point. The output tensor must be boolean. |
| `SELECT_V2` | `int8`, `int16` | no | — | `no_ctx` | The condition tensor must be boolean; the x, y and output tensors must be int8 or int16 sharing one dtype. Rank must be 8 or lower and the output shape must be the broadcast of the input shapes. |
| `WHERE` | `int8`, `int16` | no | — | `no_ctx` | The condition tensor must be int8 or int16 with rank 1 to 8. The output tensor must be int64 and must match the shape the kernel derives from the condition. |

## Convolution operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `CONV_2D` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `conditional` | Input and output must share one dtype, and float weights and bias must use that same dtype. Weight scales must be per-tensor or one per output channel on the output-channel axis (0 for OHWI). Upscale factors must be 1 or greater, and upscaling is supported on the int8 kernel path only. The weights and bias tensors must be constants so the weight sums can be precomputed. |
| `DEPTHWISE_CONV_2D` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `conditional` | Input and output must share one dtype, and float weights and bias must use that same dtype. Weight scales must be per-tensor or one per output channel on the output-channel axis (3 for 1HWO). The weights and bias tensors must be constants, with rank 3 or 4, so the weight sums can be precomputed. |
| `TRANSPOSE_CONV` | `int8`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `conditional` | Input and output must share one dtype. The int8 path requires int8 weights and, when present, an int32 bias. The float path requires weights and any bias to use the input dtype. |

## Dense operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `FULLY_CONNECTED` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `conditional` | Input and output must share one supported dtype. |

## Elementwise operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `ABS` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.31.0` | `no_ctx` | Input and output must have the same dtype and the same shape. |
| `ADD` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Both inputs and the output must share one dtype. Float ADD is elementwise only: rank-4-or-lower shapes may differ only by leading size-one dimensions; shapes above rank 4 must be identical. |
| `MAXIMUM` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `ignored` | Takes exactly two input tensors and one output tensor, all of a supported dtype. |
| `MINIMUM` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `ignored` | Takes exactly two input tensors and one output tensor, all of a supported dtype. |
| `MUL` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.32.0` | `no_ctx` | Both inputs and the output must share one dtype. Float MUL broadcasts rank-4-or-lower tensors after left-padding them to NHWC; equal shapes of any rank retain the flat kernel. Float MUL requires positive dimensions and at most INT32_MAX elements. |
| `RSQRT` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.33.0` | `no_ctx` | Input and output must share one dtype and one shape. Quantized tensors must carry a scale and a zero-point, and int16 tensors must use zero-point 0. Float RSQRT requires positive dimensions and at most INT32_MAX elements. |
| `SQRT` | `int8`, `int16`, `float16`, `float32` | `float16` | `7.33.0` | `no_ctx` | Input and output must share one dtype and one shape. Quantized tensors must carry a scale and a zero-point. float16 SQRT requires positive dimensions and at most INT32_MAX elements. |
| `SQUARED_DIFFERENCE` | `int8`, `int16` | no | — | `no_ctx` | Takes exactly two input tensors and one output tensor, all in the same integer dtype. |
| `SUB` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.32.0` | `no_ctx` | Both inputs and the output must share one dtype. Float SUB supports elementwise shapes, a scalar operand, or NumPy/TFLite broadcasting of operands of rank 4 or lower, compared in the left-padded NHWC view. Scalar-broadcast float SUB supports the NONE fused activation only. Broadcast dimensions must be positive and the element counts must fit int32. |

## Matmul operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `BATCH_MATMUL` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `conditional` | Both operands and the output must share one dtype. |

## Pooling operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `AVERAGE_POOL_2D` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `conditional` | Takes exactly one input tensor and one output tensor. Input and output must both use a supported dtype; pooling does not convert element types. |
| `MAX_POOL_2D` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `ignored` | Input and output must both use a supported dtype; pooling does not convert element types. |

## Quantization operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `DEQUANTIZE` | `int8`, `int16`, `float16`, `float32` | no | `7.35.0` | `no_ctx` | Accepts int8, int16 or float16 input and produces float32 output. Input and output shapes must match. |
| `QUANTIZE` | `int8`, `int16`, `float32` | no | `7.35.0` | `no_ctx` | Accepts float32 input to int8/int16 output, or int8-to-int8 and int16-to-int16 requantization. |

## Reduction operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `ARG_MAX` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | The output tensor must be int32. The axis attribute must be in range for the input rank. |
| `ARG_MIN` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | The output tensor must be int32. The axis attribute must be in range for the input rank. |
| `MEAN` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.32.0` | `no_ctx` | Input and output must share one supported dtype. Input rank must be 1 to 4 with positive dimensions. Reduction axes must be constant and in range; integer MEAN requires at least one axis. Float input element count must fit int32. Float output shape must match the reduction with reduced dimensions kept or removed. Integer output shape must match keep_dims; when unspecified, kept or removed dimensions are accepted. |
| `REDUCE_MAX` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.34.0` | `no_ctx` | Input and output must share one dtype; there is no requantization. |
| `REDUCE_MIN` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.34.0` | `no_ctx` | Input and output must share one dtype; there is no requantization. |
| `SUM` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.31.0` | `no_ctx` | Input and output must share one dtype; integer tensors must carry quantization parameters. The axes tensor must be constant int32 with axes in the input rank's range. Input and output rank must be at most 4. The output shape must match the reduction with reduced dimensions kept or removed. |

## Stateful operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `ASSIGN_VARIABLE` | `int8`, `int16`, `float16`, `float32` | no | `7.35.0` | `no_ctx` | Resource and assigned tensors must have the same numeric dtype. Float16 tensors require a platform with FP16 kernel support. |
| `GRU` | `float16` | `float16` | `7.35.0` | `no_ctx` | Takes two inputs (sequence and initial state) and two outputs (sequence and final state). Every tensor must be float16, and the weight tensors must be constants. Sequence tensors must be rank 3 with positive dimensions, batch-major, and batch size 1. Only reset_after=True is supported. The initial state and the final state must be distinct tensors. |
| `READ_VARIABLE` | `int8`, `int16`, `float16`, `float32` | no | `7.35.0` | `no_ctx` | Resource and output tensors must have the same numeric dtype. Float16 tensors require a platform with FP16 kernel support. |
| `SVDF` | `int8`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `required` | Kernel rank limits are 1..32767 for integer paths and 1..2147483647 for float paths. Input and output must share one dtype. On the float path the bias and weight tensors must use the input dtype. |
| `UNIDIRECTIONAL_SEQUENCE_LSTM` | `int8`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Only the standard Keras LSTM tensor subset is supported, with TANH cell activation. Projection clipping, diagonal recurrent tensors and asymmetric_quantize_inputs are not supported. Input and output must share one dtype. Weight and bias tensors must be constants whose shapes follow the input and hidden sizes. The output state and the cell state must each hold batch size times hidden size elements. On the int8 path the weights must be symmetric (zero-point 0) and the output state must be int8. On the int8 path the output state must carry the quantization of the primary output. On the int8 path the cell state must be int16 with zero-point 0 and a power-of-two scale. |

## Tensor operators

| Operator | Data types | Float dependency | Minimum ns-cmsis-nn | Context buffer | Restrictions |
| --- | --- | --- | --- | --- | --- |
| `BATCH_TO_SPACE_ND` | `int8`, `int16` | no | — | `no_ctx` | Input rank must be 3 or 4, with one block_shape entry per spatial dimension. Block shape entries must be non-zero and the batch dimension must divide by each of them. Crops must be non-negative and every computed output dimension must be positive. Input and output must share one dtype; there is no requantization. |
| `BROADCAST_TO` | `int8`, `int16` | no | — | `no_ctx` | Input and output tensors must be int8 or int16 with rank 1 to 8. The shape tensor must be a constant rank-1 int32 or int64 tensor holding only positive dimensions. The output shape must equal the shape tensor and must be a valid broadcast of the input shape. |
| `CONCATENATION` | `int8`, `int16`, `int32`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Takes between one and ten input tensors, all with the same non-zero rank as the output. Float concatenation requires 4-D NHWC tensors, and every input must use the output dtype. |
| `DEPTH_TO_SPACE` | `int8`, `int16` | no | — | `no_ctx` | Input and output must share one dtype; there is no requantization. |
| `DILATE` | `int8`, `int16`, `float16`, `float32` | `float16` | `7.35.0` | `no_ctx` | Requires DILATE options and a constant int32 dilations tensor with one entry per input dimension. Input rank must be 1 to 4 and every dilation factor must be 1 or greater. Each output dimension must equal (input dimension - 1) * dilation + 1. The padding value must be a finite constant scalar in the input dtype. Input and output must share one dtype. When both tensors carry quantization metadata, first scales must be close (numpy.isclose) and first zero-points must be equal. |
| `DYNAMIC_UPDATE_SLICE` | `int8`, `int16` | no | — | `no_ctx` | Operand, update and output tensors must be int8 or int16 and share one dtype. Operand rank must be 1 to 8, the update rank must match it, and no update dimension may exceed the operand. start_indices must be a rank-1 int32 or int64 tensor with one entry per operand dimension. The output shape must equal the operand shape. |
| `EXPAND_DIMS` | `int8`, `int16`, `int32`, `float16`, `float32` | no | `7.35.0` | `no_ctx` | Takes exactly one input tensor and one output tensor. Moves tensor bytes without numeric conversion; aliased input/output buffers require no copy. |
| `FILL` | `int8`, `int16`, `float16`, `float32` | `float16` | `7.35.0` | `no_ctx` | The fill value must use the output dtype: the emitted cast reinterprets it rather than requantizing it. |
| `GATHER` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.34.0` | `no_ctx` | The indices tensor must be int32 and the output dtype must match the data tensor. The axis must be in range for the input rank. The output rank must equal input rank plus indices rank minus batch_dims minus one. Float tensors must have positive dimensions; data rank is 1 to 4 and indices/output rank is 0 to 4. Each float-path tensor buffer must fit within INT32_MAX bytes. Float batch_dims must be nonnegative, at most axis and indices rank, and less than data rank. Float data and indices must have matching leading batch dimensions. The float output shape must match the shape inferred from data, indices, axis and batch_dims. |
| `GATHER_ND` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.34.0` | `no_ctx` | The indices tensor must be int32 and the output dtype must match the params tensor. Float option shapes must match the actual data, indices and output tensor shapes. Float tensors must have positive dimensions; data/indices rank is 1 to 4 and output rank is 0 to 4. Each float-path tensor buffer must fit within INT32_MAX bytes. Float index depth must not exceed data rank. The float output shape must equal indices.shape[:-1] followed by data.shape[index_depth:]. |
| `MIRROR_PAD` | `int8`, `int16` | no | — | `no_ctx` | Input and output tensors must be int8 or int16, rank 1 to 8, sharing one dtype, scale and zero-point. int16 tensors must use zero-point 0. The paddings tensor must have shape [rank, 2] and hold non-negative values. Paddings must stay within the input-size limits of the selected mirror mode. The output shape must equal the padded input shape. |
| `PACK` | `int8`, `int16`, `int32`, `float16`, `float32` | `float16`, `float32` | `7.33.1` | `no_ctx` | Takes at least one input, and the input count must equal both values_count and the size of the pack axis. Every input must use the output dtype and must equal the output shape with the pack axis removed. |
| `PAD` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Pre- and post-padding values must be given. Input and output must share one supported dtype. Quantized input and output must share one scale. For quantized tensors, the constant padding value must fit the output dtype range. |
| `RESHAPE` | `int8`, `int16`, `float16`, `float32` | no | `7.35.0` | `no_ctx` | Input and output must share one dtype and the same total byte count. |
| `RESIZE_BILINEAR` | `int8`, `int16` | no | — | `no_ctx` | Input and output tensors must be int8 or int16 with rank 4 or lower, sharing one scale and zero-point. The size tensor must be a constant int32 tensor of exactly two elements matching the output height and width. half_pixel_centers and align_corners must not both be true. The output batch and channel dimensions must match the input. |
| `RESIZE_NEAREST_NEIGHBOR` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.33.0` | `required` | The size tensor must be a constant int32 tensor of two elements: the kernel precomputes its coordinate map. The size tensor values must match the output height and width. Rank must be 1 to 4 with positive dimensions, and the output batch and channels must match the input. The element count and the coordinate map must both fit int32 indexing. |
| `REVERSE_SEQUENCE` | `int8`, `int16` | no | — | `no_ctx` | Input and output tensors must be int8 or int16, sharing one dtype and one shape, with rank 1 to 8. seq_lengths must be a rank-1 int32 tensor as long as the batch dimension. seq_lengths values must be non-negative and no larger than the sequence dimension. seq_dim and batch_dim must be non-negative, different, and in range for the input rank. |
| `REVERSE_V2` | `int8`, `int16`, `float16`, `float32` | no | `7.35.0` | `no_ctx` | Input and output must share one dtype and one shape, with rank 1 or higher. Axes must be unique once normalized and in range for the input rank. |
| `SCATTER_ND` | `int8`, `int16` | no | — | `no_ctx` | Updates and output tensors must be int8 or int16 and share one dtype. The indices tensor and the shape tensor must be int32, and the shape tensor must be a constant. The output shape must equal the contents of the shape tensor. The index depth must be between 1 and the output rank. Indices and updates must share their leading dimensions, and updates must match the output suffix dimensions. |
| `SHAPE` | `int8`, `int16`, `float16`, `float32` | no | `7.35.0` | `no_ctx` | None declared. |
| `SLICE` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Input rank must be 4 or lower. The output must use the input dtype. |
| `SPACE_TO_BATCH_ND` | `int8`, `int16` | no | — | `no_ctx` | Input rank must be 3 or 4, with one non-zero block_shape entry per spatial dimension. Each padded spatial dimension must divide by its block_shape entry. Input and output must share one dtype; there is no requantization. |
| `SPACE_TO_DEPTH` | `int8`, `int16` | no | — | `no_ctx` | Input and output must share one dtype; there is no requantization. |
| `SPLIT` | `int8`, `int16`, `int32`, `float16`, `float32` | `float16`, `float32` | `7.33.1` | `no_ctx` | Takes one input tensor and at least one output, and split_lengths must have one entry per output. The axis must be in range, every split length must be positive, and the lengths must sum to the axis size. Every output must use the input dtype. |
| `SQUEEZE` | `int8`, `int16`, `float16`, `float32` | `float16` | `7.35.0` | `no_ctx` | SQUEEZE only drops size-1 dimensions: input and output must share one dtype and the same byte count. |
| `STRIDED_SLICE` | `int8`, `int16`, `int32`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Input rank must be 4 or lower. The output must use the input dtype. |
| `TILE` | `int8`, `int16` | no | — | `no_ctx` | Input and output tensors must be int8 or int16 and share one dtype. The multipliers tensor must be int32, or a constant int64 tensor whose values fit int32. There must be one non-negative multiplier per input dimension. The output shape must equal the input shape multiplied elementwise by the multipliers. |
| `TRANSPOSE` | `int8`, `int16`, `float16`, `float32` | `float16`, `float32` | `7.35.0` | `no_ctx` | Input and output must both use a supported dtype. |
| `UNPACK` | `int8`, `int16`, `int32`, `float16`, `float32` | `float16`, `float32` | `7.33.1` | `no_ctx` | Takes one input tensor and at least one output, and the output count must equal both num and the axis size. Every output must use the input dtype and must equal the input shape with the unpack axis removed. |
| `ZEROS_LIKE` | `int8`, `int16`, `int32`, `float16`, `float32` | `float16` | `7.35.0` | `no_ctx` | Input and output must have the same shape and the same dtype. |

:::note[Missing an operator?]
An operator absent here is not supported by the default registry. The Python API can add custom parser/emitter registrations; see [Custom operators](https://ambiqai.github.io/helia-aot/guide/custom-operators/). Report unsupported built-in requirements with a representative model at support.aitg@ambiq.com.
:::
