AotOperator
PythonBase class for the AOT lowering of a single AIR operator.
AotOperator( op: AirOperator, model: AirModel, platform: SocPlatform, prefix: str = 'aot', attributes: dict[str, Any] | None = None,)Base class for the AOT lowering of a single AIR operator.
One instance wraps one AirOperator and drives it through the
per-operator half of the conversion: resolve() validates the operator
against the target platform and materializes any scratch or constant
tensors, plan() reports what the memory planner must reserve, and
emit() renders the operator’s C source. Subclasses implement those steps
for one TYPE and are registered into
RegistryContext.aot_operator_classes (see
register_default_aot_operators for the built-ins).
The upper-case class attributes below are the declarative contract the rest
of the pipeline reads. AotMeta rejects reassignment of an upper-case
name that a class declares in its own body, so a subclass that inherits
SUPPORTED_DTYPES without redeclaring it can still have it assigned at
runtime.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
op | AirOperator | Required | The AIR operator to wrap. |
model | AirModel | Required | The AIR model. |
platform | SocPlatform | Required | The target platform for code generation. |
prefix | str | 'aot' | Prefix for generated code files. Defaults to "aot". |
attributes | dict[str, Any] | None | None | Attributes for template values. |
TYPE
PythonTYPE: AirOpType = AirOpType.OPERATORAIR operator type this class lowers, and its registry key.
CAPABILITIES
PythonCAPABILITIES: OpCapability = OpCapability.NONEDeclared optimization contract (interning, static
reordering, shared init); see OpCapability.
USES_CMSIS_FLOAT_KERNEL
PythonUSES_CMSIS_FLOAT_KERNEL: bool = FalseWhether the float lowering dispatches to a gated ns-cmsis-nn float kernel.
CMSIS_FLOAT_KERNEL_DTYPES: frozenset = frozenset({np.float16, np.float32})Float dtypes that dependency applies to.
MIN_CMSIS_NN_VERSION_FLOAT: tuple[int, int, int] | None = NoneMinimum ns-cmsis-nn version the float
lowering requires, as (major, minor, patch), or None for the
module-wide floor.
CTX_BUF_USAGE
PythonCTX_BUF_USAGE: CtxBufUsage | None = NoneDeclared cmsis_nn_context::buf contract; see
CtxBufUsage.
CTX_BUF_BACKING_TENSOR_NAMES: tuple[str, ...] = ()SUPPORTED_DTYPES
PythonSUPPORTED_DTYPES: tuple[type, ...] | None = NoneTensor dtypes the operator accepts, or None when
the operator imposes no dtype gate of its own.
RESTRICTIONS
PythonRESTRICTIONS: tuple[str, ...] = ()User-facing sentences describing the limits validate()
enforces.
kInputTensorIndex
PythonkInputTensorIndex: int = 0kOutputTensorIndex
PythonkOutputTensorIndex: int = 0kWeightsTensorName
PythonkWeightsTensorName: str = 'weights'kMultiplierTensorName
PythonkMultiplierTensorName: str = 'multiplier'kShiftTensorName
PythonkShiftTensorName: str = 'shift'op
Pythonop: AirOperator = opmodel
Pythonmodel: AirModel = modelplatform
Pythonplatform: SocPlatform = platformprefix
Pythonprefix: str = prefixattributes
Pythonattributes: OperatorAttributes = OperatorAttributes(**attributes or {})requested_attribute_keys
Pythonrequested_attribute_keys: frozenset[str] = frozenset(attributes or {})resolved_knobs
Pythonresolved_knobs: dict[str, str] = default_knob_values()schedule
Pythonschedule: ScheduleMode = ScheduleMode.tableid
PythonReturn the operator ID.
id: strReturn the operator ID.
name
PythonReturn the operator name.
name: strReturn the operator name.
code_in_itcm
PythonWhether this operator's run code is placed in ITCM.
code_in_itcm: boolWhether this operator’s run code is placed in ITCM.
True only when code_placement is ITCM and the target has ITCM;
the operator handler warns about requests that cannot be honored.
code_placement_macro
PythonName of the <prefix>platform.h macro that places run code in ITCM.
code_placement_macro: strName of the <prefix>_platform.h macro that places run code in ITCM.
input_tensors
PythonReturn the input tensors.
input_tensors: list[AirTensor]Return the input tensors.
output_tensors
PythonReturn the output tensors.
output_tensors: list[AirTensor]Return the output tensors.
local_tensors
PythonGet all local tensors used by the operator.
local_tensors: list[AirTensor]Get all local tensors used by the operator.
QUANTIZED_FILTER_CHANNEL_AXIS: int | None = Nonehas_init
PythonWhether this operator emits a per-instance init function.
has_init: boolWhether this operator emits a per-instance _init function.
Operators whose initialization is a no-op may opt out by returning
False. When an operator opts out, it must not emit an _init
function/declaration in its templates, and the dispatch table records a
NULL init slot (the model init loop skips the call while preserving
the init callback contract). Defaults to True so existing operators
keep emitting their own init unchanged.
capabilities
PythonReturn the operator's effective optimization capabilities.
capabilities: OpCapabilityReturn the operator’s effective optimization capabilities.
The effective set is the declared :attr:CAPABILITIES widened by the
capabilities implied by the operator’s runtime behavior, so the legacy
seams remain the source of truth during migration and no operator has to
declare a capability twice:
not self.has_initimplies :attr:OpCapability.STATELESS.- A non-empty :meth:
shared_kernel_helpersor :meth:shared_activation_helpersimplies :attr:OpCapability.DESCRIPTOR_DRIVENand :attr:OpCapability.SHARED_KERNEL.
Because some seams are dtype-dependent (e.g. fully-connected only routes through a shared kernel for int8 per-channel quantization), this is an instance property rather than a class constant: it reflects the concrete resolved operator. It performs no model mutation and does not change emitted code.
direct_dispatch
PythonReturn the direct shared-kernel call target for static scheduling.
direct_dispatch: tuple[str, str] | NoneReturn the direct shared-kernel call target for static scheduling.
Under static scheduling the model schedule can call a descriptor
driven operator’s shared kernel helper directly, eliding the thin
per-node _run thunk (the main .text reclaim described in RFC
0003 Stage 4). This is only valid for operators whose _run body is
exactly return helper(ctx, &desc); – i.e. operators that route
through a single shared kernel helper with the uniform
helper(ctx, const desc_t *) signature (ADD/MUL/per-channel
int8 FULLY_CONNECTED). Shared activation helpers marshal flattened
arguments and keep mutable file-scope state, so they are intentionally
excluded and keep their _run thunk.
set_knobs
PythonSet the resolved optimization knob values before resolve().
set_knobs(values: dict[str, str]) -> NoneSet the resolved optimization knob values before resolve().
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
values | dict[str, str] | Required | Knob name to resolved value, for every knob. |
Raises
| Type | Description |
|---|---|
ValueError | When a knob is missing or unknown, or a value is ``auto`` or not one of the knob's values. |
knob_applicability
PythonWhether a knob's alternatives would change this operator's code.
knob_applicability(name: str) -> tuple[bool, str | None] | NoneWhether a knob’s alternatives would change this operator’s code.
Operators that implement a knob override this; the base operator has no knobs.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
name | str | Required | The knob's name. |
Returns
| Type | Description |
|---|---|
tuple[bool, str | None] | None | tuple[bool, str | None] | None: ``(applicable, reason when not)``, |
tuple[bool, str | None] | None | or None when the knob does not exist for this operator. |
fixed_choices
PythonChoices that change numerics or speed but are not knobs yet.
fixed_choices() -> dict[str, tuple[str, FixedSetBy]]Choices that change numerics or speed but are not knobs yet.
Returns
| Type | Description |
|---|---|
dict[str, tuple[str, FixedSetBy]] | dict[str, tuple[str, FixedSetBy]]: Choice name to ``(value, set_by)``, |
dict[str, tuple[str, FixedSetBy]] | where ``set_by`` is ``attribute``, ``default``, ``heuristic`` or |
dict[str, tuple[str, FixedSetBy]] | ``knob``. |
typed_options
PythonReturn self.op.options checked against the operator's options class.
typed_options(options_type: type[_OptionsT]) -> _OptionsTReturn self.op.options checked against the operator’s options class.
AirOperator.options is typed as the AirOperatorOptions marker
base; each AOT operator knows its concrete class and reads through
this so field access is type-checked and a mismatch (a mis-registered
parser, a hand-built graph) fails with a clear error.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
options_type | type[_OptionsT] | Required | The ``AirXxxOptions`` class this operator expects. |
Returns
| Type | Description |
|---|---|
_OptionsT | The options object, typed as ``options_type``. |
Raises
| Type | Description |
|---|---|
TypeError | If the operator carries options of a different class. |
float_io_dtypes
PythonNumpy float scalar types present on this operator's IO tensors.
float_io_dtypes() -> set[type]Numpy float scalar types present on this operator’s IO tensors.
Returns
| Type | Description |
|---|---|
set[type] | set[type]: The subset of ``{np.float16, np.float32}`` referenced by |
set[type] | the operator's input or output tensors. |
cmsis_float_dependencies
PythonFloat dtypes for which this operator needs the ns-cmsis-nn float API.
cmsis_float_dependencies() -> set[type]Float dtypes for which this operator needs the ns-cmsis-nn float API.
A returned dtype means the emitted C references the float-only
arm_nnfunctions_flt.h API (a gated arm_*_f32 / arm_*_f16
kernel or the float16_t / float32_t types it introduces) for
that precision, so the ns-cmsis-nn build must enable
ARM_NN_ENABLE_F32 / ARM_NN_ENABLE_F16.
The base implementation derives the set from
:attr:USES_CMSIS_FLOAT_KERNEL: operators that dispatch to a gated
float kernel advertise every float dtype on their IO; everything else
(integer kernels, byte movers, portable conversions) advertises none.
Operators with a per-instance dependency override this method.
Returns
| Type | Description |
|---|---|
set[type] | set[type]: Subset of ``{np.float16, np.float32}`` requiring the |
set[type] | ns-cmsis-nn float API. |
Minimum ns-cmsis-nn version this operator instance requires.
cmsis_nn_version_requirement() -> tuple[int, int, int] | NoneMinimum ns-cmsis-nn version this operator instance requires.
The base implementation applies :attr:MIN_CMSIS_NN_VERSION_FLOAT
only when the operator actually takes its float path, reusing the
same :meth:cmsis_float_dependencies contract that drives the float
header include and the ARM_NN_ENABLE_F32 / ARM_NN_ENABLE_F16
build defines. An int8 ABS therefore imposes no floor of its own while
an fp32 ABS does; a declaration only raises a module’s requirement when
it is above the repository-wide CMSIS_NN_VERSION.
Operators with an unconditional floor, or one that varies by something other than float dispatch, override this.
Returns
| Type | Description |
|---|---|
tuple[int, int, int] | None | tuple[int, int, int] | None: ``(major, minor, patch)``, or |
tuple[int, int, int] | None | ``None`` when the module-wide floor already suffices. |
kernel_entry
PythonThe ns-cmsis-nn function this operator calls, when it selects one kernel.
kernel_entry() -> str | NoneThe ns-cmsis-nn function this operator calls, when it selects one kernel.
Returns
| Type | Description |
|---|---|
str | None | str | None: The C function name, or None for an operator without a |
str | None | single selected kernel. |
get_local_tensor
PythonGet a local tensor by name.
get_local_tensor(name: str) -> AirTensorGet a local tensor by name.
tensor_zero_point
PythonReturn a tensor's quantization zero point, defaulting to 0.
tensor_zero_point(tensor: AirTensor, index: int = 0) -> intstaticmethod
Return a tensor’s quantization zero point, defaulting to 0.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
tensor | AirTensor | Required | The tensor to inspect. |
index | int | 0 | The zero-point index to read (per-tensor quant uses 0). |
Returns
| Type | Description |
|---|---|
int | The integer zero point, or 0 when the tensor is unquantized. |
intern_constant_array
PythonPromote an operator-local constant array into a model tensor.
intern_constant_array( name: str, data: np.ndarray, *, role: AirTensorKind = AirTensorKind.CONSTANT, alignment: int | None = None,) -> TensorIdPromote an operator-local constant array into a model tensor.
This routes operator metadata (e.g. per-channel quantization tables)
through the same machinery as ordinary constants: memory planning,
duplicate-constant interning (storage dedup), cold/staged residency,
and uniform tensor-pointer resolution. The tensor is registered under
op.named_tensors[name] so templates can reference it via
ctx->tensor_ptrs[...] like any other tensor.
The operation is idempotent: a previously interned tensor with the same deterministic id is replaced.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
name | str | Required | Local tensor name; also used to build the deterministic id. |
data | np.ndarray | Required | Array payload, stored as a contiguous copy. |
role | AirTensorKind | AirTensorKind.CONSTANT | Tensor kind/bucket. Defaults to ``CONSTANT``. |
alignment | int | None | None | Optional alignment hint in bytes. |
Returns
| Value | Type | Description |
|---|---|---|
TensorId | TensorId | The id of the registered tensor. |
validate
PythonValidate the operator configuration.
validate()Validate the operator configuration.
resolve
PythonPublic entry point—only runs once, even if called repeatedly.
resolve()Public entry point—only runs once, even if called repeatedly.
This method validates the AIR operator and performs any model mutations needed.
Raises
| Type | Description |
|---|---|
ValueError | If the operator declares that a kernel reads through a ``cmsis_nn_context`` buffer but allocated no backing for it (see :meth:`_assert_ctx_buf_backed`). |
on_resolve
PythonOverride this in subclasses to do the actual work.
on_resolve()Override this in subclasses to do the actual work.
ctx_buf_required
PythonWhether this concrete operator's lowering dereferences ctx->buf.
ctx_buf_required() -> boolWhether this concrete operator’s lowering dereferences ctx->buf.
Only meaningful for operators declaring
:attr:CtxBufUsage.CONDITIONAL, which MUST override this. Evaluate it
against the same inputs the operator’s scratch sizing uses (dtype,
options, platform capability) so the two can never disagree: the guard
exists precisely to catch the case where sizing says “0 bytes” and the
dispatched kernel says “I read that buffer”.
Returns
| Type | Description |
|---|---|
bool | ``True`` if at least one kernel this operator can dispatch reads or |
bool | writes through ``ctx->buf``. |
Raises
| Type | Description |
|---|---|
NotImplementedError | If called on an operator that did not declare :attr:`CtxBufUsage.CONDITIONAL` and override this method. |
plan
PythonOperator planning step.
plan()Operator planning step.
shared_kernel_helpers
PythonReturn the shared operator-kernel helpers this operator requires.
shared_kernel_helpers() -> list[KernelHelper]Return the shared operator-kernel helpers this operator requires.
This mirrors :meth:shared_activation_helpers but for general operator
kernels (e.g. elementwise ADD/MUL or FULLY_CONNECTED). An
operator that routes its _run body through a shared runtime helper
(instead of an inlined per-node call site) declares the helper
descriptors here so the code generator emits each distinct helper
exactly once. Helpers are keyed by signature, so distinct operator
variants map to distinct helpers and never collide. The base
implementation requires no shared kernels.
Returns
| Type | Description |
|---|---|
list[KernelHelper] | A list of :class:`KernelHelper` descriptors (see |
list[KernelHelper] | mod:`helia_aot.aot.operators.kernel_runtime`). Empty for operators |
list[KernelHelper] | that do not use a shared kernel helper. |
Return the shared LUT-activation helpers this operator requires.
shared_activation_helpers() -> list[dict[str, object]]Return the shared LUT-activation helpers this operator requires.
Operators that emit their hot loop via a shared runtime helper (rather than an inlined per-node body) declare the helper descriptors here so the code generator can emit each distinct helper exactly once. The base implementation requires no shared helpers.
Returns
| Type | Description |
|---|---|
list[dict[str, object]] | A list of helper descriptor dicts (see |
list[dict[str, object]] | func:`helia_aot.aot.operators.activation_runtime.activation_lut_helper`). |
list[dict[str, object]] | Empty for operators that do not use a shared activation helper. |
compute_values
PythonCompute the values for the operator source code template.
compute_values() -> dict[str, Any]Compute the values for the operator source code template.
Args:
Returns
| Type | Description |
|---|---|
dict[str, Any] | dict[str, Any]: Dictionary of values for the operator template. |
print_info
PythonDebug print the template values for the operator.
print_info(verbose: int = 0)Debug print the template values for the operator.
emit
PythonGenerate the source code for the operator.
emit(save_path: Path)Generate the source code for the operator.
This method should be overridden by subclasses.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
save_path | Path | Required | Path to save the generated source code. |