Skip to content
heliaAOT
HELIA HUB

Configuring a conversion

One configuration object, ConvertArgs, describes a conversion. You can write it as YAML, pass it as command-line flags, or build it in Python. All three reach the same validated object, and every field is documented in the Configuration reference.

Use a model that already converts on your chosen target. Keep the model path, target and output directory in one file before changing individual controls:

convert.yaml
model:
path: ./model.tflite
module:
path: ./out
name: my_model
platform:
name: apollo510_evb
memory:
dump_residency_json: true
Terminal window
helia-aot convert --path convert.yaml

A successful conversion writes out/my_model/, including the generated source, a README describing the model and its integration, and aot_residency.json. Keep the YAML file and any explicit CLI overrides with this output. Review configuration warnings and the placement report before building the firmware. If that module already exists, select a new output directory or use --force deliberately. Conversion success does not establish numerical correctness; follow Testing and automation before tuning.

A YAML file is passed with --path and must be a mapping. The command line is deep-merged over it, and only flags you actually typed take part: every option defaults to nothing rather than to the field’s default, so naming a flag is what overrides YAML, not accepting a default.

Terminal window
helia-aot convert --path convert.yaml --module.path ./out --force

The merge is per key, not per block, so a command-line --platform.memories entry replaces that one memory and leaves the rest of the YAML block standing. Nested fields become dotted, kebab-cased flags (--memory.auto-hydrate-constants), booleans become a pair (--test.enabled and --no-test.enabled), and list options accept repeated values, space-separated values, or a bare flag meaning an explicit empty list.

In Python the equivalent is ConvertArgs.from_yaml(path), then AotConverter(config).convert().

Three sources can supply a value, and they are consulted in one order: a flag you typed, then the YAML file, then the field’s own default. Take this file:

convert.yaml
model:
path: kws_ref.tflite
module:
path: ./out
type: zephyr

and this command:

Terminal window
helia-aot convert --path convert.yaml --module.type cmake
Setting Flag YAML Default Value used Why
module.type cmake zephyr neuralspot cmake A flag you typed beats the file.
module.path not given ./out output.zip ./out The file beats the default.
module.prefix not given not given aot aot Nothing overrode the default.
module.schedule not given not given table table The same, and the reason a flag you did not type cannot silently reimpose a default.

Omitted flags do not replace YAML values. A wrapper script can override one setting while retaining the rest of the saved configuration.

The top level is strict: an unknown key in ConvertArgs is rejected. Nested blocks are warn-first: an unknown key produces a deprecation warning with a did-you-mean suggestion and is dropped, which will become a hard error in a later release. Treat those warnings as failures.

Two keys are already hard errors because their behaviour is gone: memory.persistent_storage and memory.constant_residency.

A failed validation exits with code 2 and one message per field. A conversion that starts and then fails exits 1.

Block What it sets Reference
model Input path, subgraph index, name, description, version ModelArgs
module Output path, module type, name, prefix, schedule mode ModuleArgs
test Generated test case: tolerance, golden data, iterations, state feedback TestArgs
transforms Which graph rewrites run, and with what options TransformSpec
memory Planner, constraints, tensor rules, arena ownership, residency report MemoryArgs
platform Registered target name, or a complete custom target definition PlatformArgs
operators Attribute rules and optimization knob overrides applied to operators OperatorRuleset
optimization Goal, approximation gate and knobs; see Performance and accuracy options OptimizationArgs
documentation The experimental offline HTML site DocumentationArgs
top level verbose, force, log_file ConvertArgs

Attribute rules tune how the model is compiled and laid out without changing the graph. They come in two lists: operators at the top level, and memory.tensors.

A rule has three keys:

  • type - the entity type. For operators, an operator name such as CONV_2D; for tensors, a lowercase kind: constant, persistent or scratch. * matches everything.
  • id - one entity id, or a list of them, as strings. Use operator ids from the resolved graph and tensor ids from its residency report.
  • attributes - the settings to apply.

All matching rules are collected, sorted by specificity, then merged field by field in that order, so a later rule wins. Specificity, lowest first:

  1. neither type nor id (or type: "*" with no id), the catch-all
  2. id only
  3. type only
  4. both type and id, the most specific

Two rules of equal specificity are settled by position: the later one in the file wins. Run with --verbose 2 to see the attributes each operator ended up with.

A type-only rule outranks an id-only rule even when the id-only rule appears later. To make a per-node exception to a type rule, specify both:

operators:
- type: "*"
attributes: { code_placement: MRAM }
- type: CONV_2D
attributes: { code_placement: ITCM }
- id: "1"
attributes: { code_placement: MRAM }
- type: CONV_2D
id: "1"
attributes: { code_placement: MRAM }

For a CONV_2D node whose id is "1", the last rule restores MRAM code placement. Without that last rule, the type rule wins and code goes to ITCM. Rules merge individual attributes; they do not replace the whole attribute dictionary.

code_placement controls the generated operator run code. MRAM keeps the default placement; ITCM requests the target’s ITCM section. Initialization and the model schedule keep their default placement, and external kernel-library functions follow your linker script. See ITCM placement for module-specific macros, memory budgets and verification.

operators:
- type: "*"
attributes:
code_placement: MRAM
- type: FULLY_CONNECTED
attributes:
code_placement: ITCM

An operator rule can also carry optimization, which overrides the model-wide optimization knobs for the operators it matches, with the same precedence as attributes. The accumulation knob, for example, selects packed FP16 weights (fast) or the standard layout (precise) for FP16 1x1 convolutions and fully connected layers. See Performance and accuracy options for every knob, the goal and the approximation gate.

optimization:
accumulation: fast
operators:
- type: CONV_2D
optimization:
accumulation: precise

scratch_placement is not implemented and produces a warning. To place scratch buffers, use memory.tensors rules with type: scratch and an appropriate memory attribute. Operator code placement does not move model tensors.

Some operators take attributes of their own through the same mechanism - RSQRT takes lut_mode, HARD_SWISH takes use_lut and compat_variant, ETHOS_U takes require_platform_npu. The operator catalog is where per-operator attributes are listed.

Tensor rules control placement: memory selects the memory a tensor is read from, and constant_destination_memory selects a different runtime memory to stage a constant into.

memory:
tensors:
- type: constant
attributes:
memory: MRAM
- type: scratch
attributes:
memory: DTCM

Memory covers what the placements mean and when staging is worth it.

For a registered target such as apollo510_evb, set platform.name and use its declared capabilities and memory map. platform.memories replaces the sizes it names for that conversion; other platform fields do not override a registered target, and the converter warns and ignores them. To limit the memory available to this model, use platform.memories or memory.constraints.

An unregistered name requests a conversion-local custom target. It requires cpu, speeds, memories, preferred_memory_order and min_alignment; set capabilities to what the hardware actually supports. Start from the Targets reference and review the whole definition. Declaring a capability cannot add it to the silicon. Target resolution occurs during conversion, after configuration construction.

Leave transforms absent for the built-in pipeline. All four rewrites are enabled and run in this order, using these exact registry names:

Name Purpose
FOLD_STATIC_SHAPE_EXPRESSIONS Fold supported shape-building expressions whose values are known at conversion time.
DEPTHWISE_TO_CONV Rewrite eligible single-input-channel depthwise convolutions as convolutions.
PRUNE_IDENTITY_OPS Remove supported operations that leave their inputs unchanged.
TRANSPOSE_REVERSE_CONV Rewrite eligible integer transpose convolutions as convolutions.

To isolate the effect of a rewrite, disable it explicitly, then reconvert, validate the outputs and compare the emitted module:

transforms:
- name: PRUNE_IDENTITY_OPS
enabled: false
options: {}

Names are case-sensitive; Python class names such as PruneIdentityOps are not registry names. A wildcard entry (name: "*") sets the default enabled state, and named entries override it. Changing list order does not reorder the pipeline. Put options on a named entry; wildcard options are not forwarded to every transform.

For example, DEPTHWISE_TO_CONV accepts output_channel_threshold (default 1) and TRANSPOSE_REVERSE_CONV accepts input_channel_threshold (default 16). A candidate must also satisfy the rewrite’s shape and dtype conditions. These thresholds control eligibility, not an assurance of faster inference.

module.schedule picks how _model_run invokes each operator.

  • table - a function-pointer table walked by a loop, with the per-node callback seam intact. Choose this when an RTOS or a profiler needs to observe or interpose on each node. This is the default.
  • static - a straight-line sequence of direct calls. No table, no loop, no indirect dispatch, and no callback seam. Choose this when you want the compiler to see the whole sequence.

Both modes emit the same public API, the same status codes, the same lifecycle latch and the same constant hydration. Only the dispatch differs.

Every conversion writes a LICENSE and a README.md into the module root. The README describes model I/O, arenas, operator distribution, library requirements and integration. It does not contain the conversion configuration as YAML; retain your configuration file and explicit CLI overrides separately.

documentation.html adds an offline HTML site to that:

documentation:
html: true

It builds a small MkDocs site into a docs/ directory in the module root, with a page each for the model, the memory plan, usage, configuration and the licence, plus a Mermaid rendering of the graph. The configuration page serializes the validated configuration values. It is not an independent check of effective hardware settings: configured fields ignored during registered target resolution can still appear there. Only the built site is left behind; the Markdown sources and the MkDocs config it used are removed after the build.

The field is marked experimental in DocumentationArgs and defaults to false. MkDocs is imported only when the flag is set, so leaving it off costs nothing.

verbose accepts 0 to 3. Levels 2 and 3 both enable debug logging; level 3 also expands AIR operator details. The Results summary remains visible at every level.

Level What you see
0 Errors and the Results summary; progress is suppressed.
1 Stage milestones. The default.
2 Debug details, including resolved attributes and planner diagnostics.
3 Debug output plus expanded AIR tensor and option details.

Set log_file to also write the log to a file. Durable results print to standard output; progress, logs and errors go to standard error, so a script can capture one without the other. Validation warnings raised while the configuration is being loaded are buffered and replayed once logging is set up, so they are never lost to ordering.