Skip to content
heliaAOT
HELIA HUB

Release notes

This page is rendered from the repository CHANGELOG.md, which release-please writes from the conventional-commit type of every squashed pull request. Releases are newest first. The 6 most recent releases are here in full; the whole history is in CHANGELOG.md.

Unreleased — Tensor Packaging (breaking)

Section titled “Unreleased — Tensor Packaging (breaking)”

This branch consolidates per-tensor placement under a single memory.tensors: rule surface and reshapes the runtime contract for caller-supplied arenas. Two YAML keys and several C-side symbols are removed; downstream consumers should migrate before upgrading.

MemoryArgs now rejects these legacy keys with an actionable error naming the new equivalent:

  • memory.persistent_storage → use a per-tensor rule under memory.tensors, e.g.
    memory:
    tensors:
    - type: persistent
    memory: SRAM
  • memory.constant_residency → use per-tensor rules under memory.tensors with memory: (source) and optional constant_destination_memory: (staged runtime memory).
  • Removed: per-tensor <prefix>_tensor_<name>_size macros (e.g. <prefix>_tensor_0_size, <prefix>_tensor_<op>_<role>_size). Per-tensor byte size now lives on the descriptor at <prefix>_tensor_descriptors[i].size. Public _size macros are still emitted for model inputs/outputs, where they are part of the stable I/O API.
  • Changed: <prefix>_tensor_descriptor_t no longer carries a data_ptr field. Every tensor (constant, persistent, or scratch) is bound to an arena slot; the kernel-visible address is <prefix>_arena_buffers[descriptor.region] + descriptor.offset. Downstream code that read descriptor.data_ptr directly must switch to that formula.

Generated C-API changes (memory.allocate_arenas: false)

Section titled “Generated C-API changes (memory.allocate_arenas: false)”
  • <prefix>_bind_arena(region, buffer, size) — bind state is module-global; no ctx pointer parameter.
  • <prefix>_bind_arenas(buffers, sizes, n) — same; n must equal <prefix>_num_arena_buffers.
  • Removed: <prefix>_arena_table_entry_t, <prefix>_arena_buffers_table_count. Use <prefix>_arena_sizes[<prefix>_num_arena_buffers] for per-region capacity introspection.
  • <prefix>_hydrate_constants(ctx) now takes a non-const <prefix>_model_context_t *.
  • Every region must be bound before <prefix>_model_init, which internally invokes <prefix>_context_init and returns non-zero if any region is unbound.
  • memory.auto_hydrate_constants — deprecated, runtime no-op. <prefix>_model_init now always invokes <prefix>_hydrate_constants between <prefix>_context_init and the operator init loop, regardless of this flag. This closes a race where an operator’s _init hook could read constants (e.g. arm_convolve_weight_sum reading weights) before manual hydration was driven by the application. The flag is retained for backwards compatibility with existing YAML configs but no longer changes generated runtime behavior. Callers that need a custom hydration mechanism (DMA, async pre-stage, decompression, model swap) override the weak <prefix>_hydrate_constants symbol — model_init invokes the override at the same fixed point. The default helper is idempotent so pre-hydrating from the caller before model_init is safe. Setting the flag to false emits a soft WARNING at convert time so users still toggling the legacy YAML knob know the runtime no longer honors it. <prefix>_model_run returns status 200 only as a defense-in-depth check (e.g. when <prefix>_clear_hydrated() was called after a successful model_init without re-running it). The latch is observable via <prefix>_is_hydrated() and resettable via <prefix>_clear_hydrated() (also called automatically by every context_init).
  • <prefix>_bind_arena() now rejects misaligned caller buffers with status 4. The required alignment is exposed at <prefix>_arena_alignments[<prefix>_num_arena_buffers] (power-of-two per region, matches the alignas(...) applied to internal-arena builds).
  • memory.dump_residency_json (default false) writes <prefix>_residency.json with schema_version: 3 (top-level plan_hash, top-level tensor_layout_hash, per-arena region_id). The two hashes have disjoint scope: plan_hash fingerprints the arena envelope only (per-region role / memory / source_memory / size / alignment / is_staged) and mirrors the generated <PREFIX>_PLAN_HASH C macro for arena-ABI drift detection across separately-compiled binaries; tensor_layout_hash fingerprints per-tensor placement (tensor_id / role / memory / offset / size) and captures drift the envelope hash misses (two tensors swapping offsets inside the same arena, a tensor migrating between arenas of identical shape, etc.). Honors --force for overwrites.

See tensor-packaging how-to for a worked migration.

ty runs at its default rule severities and blocks the commit, and CI, on any error-level diagnostic; no rule is demoted and a unit test pins the warning count at zero. Clearing the backlog changed a few observable behaviors:

  • AirInterpreter.set_input and AirInterpreter.get_output are keyword-only. In-tree callers already passed key=/wrap=, but an external caller doing get_output(0) must add the keyword.
  • AirTensor.shape, .dtype and .ctype raise ValueError when the value is unset or the dtype has no C mapping, instead of returning None. .ctype previously let a None reach the templates, where it rendered as the literal None in generated C.
  • AirTensor.quant and AirQuantizationParameters.scales / .zero_points are the checked way to read quantization; they raise a ValueError naming the tensor when it carries none.
  • AotOperator.typed_options() checks the operator’s options against its concrete AirXxxOptions class and raises TypeError on a mismatch.
  • Reading a LiteRT flatbuffer field that the object API left unset now raises a ValueError naming the field rather than failing later on None.
  • compute_tensor_ctype rejects STRING tensors instead of returning a -1 sentinel.
  • The test floor moved to pytest>=9.1.1; pytest 8’s skip/fail wrappers are not statically analyzable.
  • add a knob registry and split the optimization plan and report (#521) (c7dba00), closes #517
  • add the optimization section, the accumulation knob and a per-layer optimization plan (#516) (57a80e7), closes #513
  • emit NT_N_PACKED weights for float 1x1 convolution and fully connected (#508) (e119e23), closes #507
  • bump the docs reference exports with each release (#505) (58c2aac), closes #504
  • infer reduction shapes from explicit AIR metadata (#509) (d62223e), closes #479
  • keep the docs exports stable under release-please’s version bump (#526) (1a0874f), closes #525
  • restore neutral active top navigation (#514) (8d861db)
  • migrate site navigation, guides and reference content (#473) (f1a3ea6)
  • polish AOT hero and product navigation (#518) (70e3edc)
  • generated modules, integer-only ones included, no longer build against ns-cmsis-nn v7.32.x to v7.34.x.
  • add greedy-by-size and hill-climb experimental memory planners (#359) (6ac4d16)
  • declare the public Python API with all and an API manifest (d2661c3), closes #455
  • dispatch float PACK, UNPACK and SPLIT to native ns-cmsis-nn kernels (#429) (20f407b)
  • export the configuration, CLI, target, error and module-layout reference data (b252649)
  • place operator code in ITCM and document ITCM tensor placement (#493) (849b4ab)
  • resolve shape expressions in the shape propagation fixed point (#476) (e537d29)
  • route 1D dilated depthwise conv to the optimized ns-cmsis-nn kernels (#503) (49a5ee4), closes #502
  • support FP16 and FP32 nearest-neighbor resize (#404) (191bed4)
  • accept JSON values for structured list flags (#484) (0c5a6f0), closes #481
  • accept scalar reduce axes and honor keepDims for integer extrema (#446) (fa0a562)
  • carry SUM keepDims and require the exact output shape (#475) (6217d3e)
  • drop the obsolete MAX/MIN/CLAMP undef prologue from the Zephyr test case (#478) (fd074a4), closes #305
  • fail fast on out-of-range tensor dims in the buffer sizers (#477) (2db5429)
  • infer dynamic shapes and fold static shape expressions (#439) (81f0076)
  • keep emitted operator sources clean under -Wunused-parameter (#491) (ddeebcd), closes #406
  • match attribute rule types case-insensitively (#482) (8077a0d), closes #474
  • preserve INT16 hard swish prescale precision (#468) (ec917d1)
  • preserve INT8 HARD_SWISH precision when deriving prescale (d68e851)
  • raise on undefined template references during codegen (#423) (494aa30)
  • remove the conversion work directory when the conversion ends (#483) (2a9be27), closes #467
  • require ns-cmsis-nn v7.35.0 for every generated module (#498) (76e6961), closes #496
  • scan registry discovery namespaces to a fixed point (#489) (9eaecda), closes #488
  • tidy conversion edge cases around memory constraints, work dirs and golden data (#495) (ba3acc8), closes #485 #486 #487 #490
  • warn when SVDF per-channel quantization is ignored (#494) (8502692), closes #329
  • consolidate strict template contract guidance (798a475)
  • generate the Python, configuration, CLI, target and module reference (21e2068)
  • scaffold the Astro site in docs/ and relocate MkDocs sources to mkdocs/ (6f55939), closes #454
  • add explicit input and output byte-count macros (#426) (c0926c3)
  • add native float argmin and argmax support (bf0fd47)
  • add native float gather and extrema support (d5cb054)
  • adopt native FP16 and FP32 RSQRT kernels (#425) (ce6c433)
  • adopt native FP16 SQRT from CORE 7.33 (#421) (83c9f3d)
  • lower non-scalar float SUB broadcasts to native kernels (0820177)
  • support native float HARD_SWISH (#410) (20ea2ce)
  • declare FP16 utility type dependencies (#422) (958432e)
  • expand per-tensor convolution quantization to every output channel (#411) (1375535)
  • finish typed parser errors and quantization hints (8d65c66)
  • honor precise float requirements in CMake and NSX (e546fc7)
  • isolate generated layer parameters across model modules (#408) (6140aff)
  • preserve integer mean logical output shapes (edd88be)
  • reject incompatible integer fully connected filters (a37f767)
  • reject int16 FULLY_CONNECTED filters no int16 kernel can honour (#418) (3c35331)
  • reject invalid concatenation axes before code generation (9ea546a)
  • lower native float MEAN and broadcast MUL (19d4d34)
  • lower native float MEAN and broadcast MUL (19d4d34)
  • preserve scalar constant rank when parsing LiteRT models (#402) (54e4217)
  • generated modules require ns-cmsis-nn v7.32.0 or newer; the generated common header fails to compile against older releases.
  • accept float16/float32 in SQUEEZE, FILL, ZEROS_LIKE, and DILATE with an operator-named dtype error (#353) (3520a96)
  • add float16 rolled and unrolled GRU support (4c3e974)
  • e2e: report passing tolerance margins (#333) (259c763)
  • wire float16/float32 SUB, STRIDED_SLICE, SLICE, and float16 SPLIT to the ns-cmsis-nn float kernels (#352) (d15ac41)
  • align generated code with ns-cmsis-nn 7.32.0 contracts (#386) (edc5176)
  • avoid shared operator attribute defaults (f42a245)
  • declare FP16 on every Cortex-M55 target so list-targets matches the supports_fp16 gate (#351) (bc2a2ea)
  • e2e: migrate to CMSIS 6 and Cortex DFP (#332) (a10be71)
  • keep config-derived custom platforms run-local instead of registering them globally (#334) (188660b)
  • pin ns-cmsis-nn at v7.32.0 and size the float depthwise scratch like its sizer (#394) (e2e10f7)
  • raise ConfigValueError with a hint for every user-facing operator validation failure (#400) (b2bb81e)
  • reject float16 and float32 on the six operators ns-cmsis-nn ships integer kernels for (#396) (14227a8)
  • report aot_tensor_io_t.size in bytes for non-int8 I/O tensors (#380) (9726d4e)
  • the at110 platform name is now atomiq110; --platform.name at110 no longer resolves. at110 shipped in v0.18.0, but Atomiq is not in production and nothing downstream has pinned the name yet.
  • add atomiq110 board support and gate Ethos-U dispatch on NPU capability (#276) (8cee4b5)
  • add broadcast_to op support (bc539f2)
  • add broadcast_to op support (bc539f2)
  • add dynamic_update_slice op support (13980ed)
  • add dynamic_update_slice op support (13980ed)
  • add mirror_pad op support (b46bb4b)
  • add mirror_pad op support (b46bb4b)
  • add RESIZE_BILINEAR and DILATE operators for int8/int16x8 (f052d51)
  • add reverse_sequence op support (68e2d6b)
  • add reverse_sequence op support (68e2d6b)
  • add scatter_nd op support (70b9e40)
  • add scatter_nd op support (70b9e40)
  • add select_v2 op support (c9ad0eb)
  • add select_v2 op support (c9ad0eb)
  • add tile op support (1570971)
  • add tile op support (1570971)
  • add where op support (6ec8a75)
  • add where op support (6ec8a75)
  • aot: add float kernel support across operator pipeline (#246) (73d9138)
  • carry recurrent state through the float LSTM path (a41df71)
  • carry recurrent state through the unidirectional sequence LSTM path (f3777b6)
  • carry recurrent state through the unidirectional sequence LSTM path (f3777b6)
  • ci: file an issue when the weekly release-model E2E run fails (#299) (3fbe40b)
  • cli: console presentation layer for conversion runs (#264) (1066679)
  • cli: help panels, factory defaults, quick-start epilog, banner (#261) (c263d26)
  • config: warn on unknown nested config keys with did-you-mean (#268) (6363014)
  • declare apollo510l compatibility in nsx module manifest (6c4fb78)
  • document and expose supported HeliaAOT target names (3b5880a)
  • document and expose supported HeliaAOT target names (3b5880a)
  • expose supported target names via CLI and add M55 targets (fc570ba)
  • float: enable FP16/FP32 for ABS, PRELU, and SUM via ns-cmsis-nn float kernels (#279) (2be2bd5)
  • gate releases on manifest-backed model-corpus E2E (#277) (0c41e19)
  • guard resolve against NULL ctx.buf in kernels that dereference it (#325) (9abc3b0), closes #316
  • rename Apollo330P target and match target names case-insensitively (3cb2215)
  • support Python 3.13 and 3.14 (8644efa)
  • support stateful unidirectional sequence LSTM (b0449a1)
  • tensor attribute-key warnings and programmatic API surface (#270) (370d954)
  • typed errors with hints and clean CLI exit codes (#263) (7f031fa)
  • Add missing doxygen blocks to header (8700eb0)
  • add, mul, sub, expand_dims, max, min, gather bugs (dbd7a27)
  • address Copilot review feedback on shape propagation (bd97015)
  • adopt context-only run signature in dilate and resize_bilinear templates (421340c)
  • align svdf recurrent state validation (#290) (bda8750)
  • aot: narrow extern “C” guard scope in generated headers (#250) (a5154f1)
  • block unresolved dims in shape rules and guard the memory planner (0d274ba)
  • build the CLI on Python 3.14 (c29139c)
  • Correct broadcast_to and batch_to_space rules with proper conditions (8a0b054)
  • count rule-confirmed shapes as resolved and claim rank-4-only STB/BTS (3208948)
  • e2e: clone the public ns-cmsis-nn without a credential (#291) (a22304e)
  • e2e: seed Keras initializers so bidirectional-GRU weights are reproducible (#314) (116f176), closes #301
  • emit scalar gather indices as rank-1 for heliaCore kernel contract (3399a8d)
  • fix add, mul, sub, expand_dims, max, min, gather bugs (dbd7a27)
  • fix add, mul, sub, expand_dims, max, min, gather bugs (4f41a11)
  • Handle potential edge case where input batch shape is also negative (not just shape signature) (90d1021)
  • harden rank-0 and expand_dims emission found by adversarial review (7611299)
  • harden RESIZE_BILINEAR/DILATE validation, correct rounding, drop broken MVE path (95b4199)
  • Improve Windows compatibility (d0cffaa)
  • keep includes outside extern “C” guard in new header templates (1d7e2d4)
  • make e2e stimulus deterministic (#288) (101aec9)
  • make generated Zephyr test case CMSIS 6 compatible (#287) (0cc6e13)
  • make shape propagation rule ownership explicit (722b642)
  • make shape propagation rule ownership explicit (722b642)
  • name custom-platform fields in both YAML and CLI spellings (160e7ab)
  • normalize negative concat axis before the ownership test (e7cf8ff)
  • raise the ns-cmsis-nn floor to 7.31.0 and pin e2e to match (#340) (59f4c6f), closes #304
  • Reject index depth == 0 (114a85b)
  • reject SVDF weights_feature/bias dtypes the kernels cannot consume (#312) (90c661e)
  • release: publish the GitHub Release only after model-corpus qualification (#298) (85eda59)
  • Remove _PUT_IN_MRAM_INIT to const mapping in test_emit-stage.py (a765e04)
  • remove const from MRAM init placement macro (3e3d519)
  • remove const from MRAM init placement macro (3e3d519)
  • remove const from MRAM init placement macro (dafd5b4)
  • remove duplicate numpy import in parser tests (a8dc4a7)
  • replace in-place ndarray shape assignment (6edada0)
  • resolve merge conflicts with main (e14c183)
  • resolve merge conflicts with main (305d13f)
  • restore CMSIS-NN no-clip sentinel and require persistent LSTM state (2641268)
  • restore scalar-scalar support for maximum and minimum (0a54eb0)
  • size int8 batch-matmul scratch from the RHS row count (#296) (28636e9)
  • size transpose-conv int8 scratch so the kernel gets a real buffer (#313) (bfe683f)
  • update fp16 golden test for the renamed LSTM e2e generator (9bdf93b)
  • validate SPLIT output shapes against the split lengths before code generation (#344) (44bd27c), closes #322
  • validate SUM output shape, rank, and constant axis against the kernel contract (#326) (b82b256), closes #280 #207
  • validate SVDF shapes against the ns-cmsis-nn kernel contract (#324) (9e3535d), closes #317 #280
  • write generated files as UTF-8 with LF newlines (4508272)
  • document exact-rational vs TFLite 10-bit divergence bounds (a6ab417)
  • drop interpreter pin from pipx example (3c87204)
  • ethos-u: correct driver contract and document known gaps (#266) (364cf21)
  • tighten recurrent-state wording and pass byte length to arm_memset_s8 (4d30064)
  • wire agent instructions into every tool and add operational guidance (#272) (c68882a)

Earlier releases are in CHANGELOG.md.