Release notes
This page is rendered from the repository CHANGELOG.md, which release-please writes from the conventional-commit type of every squashed pull request. Releases are newest first. The 6 most recent releases are here in full; the whole history is in CHANGELOG.md.
Unreleased — Tensor Packaging (breaking)
Section titled “Unreleased — Tensor Packaging (breaking)”This branch consolidates per-tensor placement under a single
memory.tensors: rule surface and reshapes the runtime contract for
caller-supplied arenas. Two YAML keys and several C-side symbols are
removed; downstream consumers should migrate before upgrading.
Removed YAML keys
Section titled “Removed YAML keys”MemoryArgs now rejects these legacy keys with an actionable error
naming the new equivalent:
memory.persistent_storage→ use a per-tensor rule undermemory.tensors, e.g.memory:tensors:- type: persistentmemory: SRAMmemory.constant_residency→ use per-tensor rules undermemory.tensorswithmemory:(source) and optionalconstant_destination_memory:(staged runtime memory).
Generated C-API changes (all build modes)
Section titled “Generated C-API changes (all build modes)”- Removed: per-tensor
<prefix>_tensor_<name>_sizemacros (e.g.<prefix>_tensor_0_size,<prefix>_tensor_<op>_<role>_size). Per-tensor byte size now lives on the descriptor at<prefix>_tensor_descriptors[i].size. Public_sizemacros are still emitted for model inputs/outputs, where they are part of the stable I/O API. - Changed:
<prefix>_tensor_descriptor_tno longer carries adata_ptrfield. Every tensor (constant, persistent, or scratch) is bound to an arena slot; the kernel-visible address is<prefix>_arena_buffers[descriptor.region] + descriptor.offset. Downstream code that readdescriptor.data_ptrdirectly must switch to that formula.
Generated C-API changes (memory.allocate_arenas: false)
Section titled “Generated C-API changes (memory.allocate_arenas: false)”<prefix>_bind_arena(region, buffer, size)— bind state is module-global; noctxpointer parameter.<prefix>_bind_arenas(buffers, sizes, n)— same;nmust equal<prefix>_num_arena_buffers.- Removed:
<prefix>_arena_table_entry_t,<prefix>_arena_buffers_table_count. Use<prefix>_arena_sizes[<prefix>_num_arena_buffers]for per-region capacity introspection. <prefix>_hydrate_constants(ctx)now takes a non-const<prefix>_model_context_t *.- Every region must be bound before
<prefix>_model_init, which internally invokes<prefix>_context_initand returns non-zero if any region is unbound.
New / renamed knobs
Section titled “New / renamed knobs”memory.auto_hydrate_constants— deprecated, runtime no-op.<prefix>_model_initnow always invokes<prefix>_hydrate_constantsbetween<prefix>_context_initand the operator init loop, regardless of this flag. This closes a race where an operator’s_inithook could read constants (e.g.arm_convolve_weight_sumreading weights) before manual hydration was driven by the application. The flag is retained for backwards compatibility with existing YAML configs but no longer changes generated runtime behavior. Callers that need a custom hydration mechanism (DMA, async pre-stage, decompression, model swap) override the weak<prefix>_hydrate_constantssymbol —model_initinvokes the override at the same fixed point. The default helper is idempotent so pre-hydrating from the caller beforemodel_initis safe. Setting the flag tofalseemits a softWARNINGat convert time so users still toggling the legacy YAML knob know the runtime no longer honors it.<prefix>_model_runreturns status200only as a defense-in-depth check (e.g. when<prefix>_clear_hydrated()was called after a successfulmodel_initwithout re-running it). The latch is observable via<prefix>_is_hydrated()and resettable via<prefix>_clear_hydrated()(also called automatically by everycontext_init).<prefix>_bind_arena()now rejects misaligned caller buffers with status4. The required alignment is exposed at<prefix>_arena_alignments[<prefix>_num_arena_buffers](power-of-two per region, matches thealignas(...)applied to internal-arena builds).memory.dump_residency_json(defaultfalse) writes<prefix>_residency.jsonwithschema_version: 3(top-levelplan_hash, top-leveltensor_layout_hash, per-arenaregion_id). The two hashes have disjoint scope:plan_hashfingerprints the arena envelope only (per-region role / memory / source_memory / size / alignment / is_staged) and mirrors the generated<PREFIX>_PLAN_HASHC macro for arena-ABI drift detection across separately-compiled binaries;tensor_layout_hashfingerprints per-tensor placement (tensor_id / role / memory / offset / size) and captures drift the envelope hash misses (two tensors swapping offsets inside the same arena, a tensor migrating between arenas of identical shape, etc.). Honors--forcefor overwrites.
See tensor-packaging how-to for a worked migration.
Type checking is now a blocking gate
Section titled “Type checking is now a blocking gate”ty runs at its default rule severities and blocks the commit, and CI,
on any error-level diagnostic; no rule is demoted and a unit test pins
the warning count at zero. Clearing the backlog changed a few observable
behaviors:
AirInterpreter.set_inputandAirInterpreter.get_outputare keyword-only. In-tree callers already passedkey=/wrap=, but an external caller doingget_output(0)must add the keyword.AirTensor.shape,.dtypeand.ctyperaiseValueErrorwhen the value is unset or the dtype has no C mapping, instead of returningNone..ctypepreviously let aNonereach the templates, where it rendered as the literalNonein generated C.AirTensor.quantandAirQuantizationParameters.scales/.zero_pointsare the checked way to read quantization; they raise aValueErrornaming the tensor when it carries none.AotOperator.typed_options()checks the operator’s options against its concreteAirXxxOptionsclass and raisesTypeErroron a mismatch.- Reading a LiteRT flatbuffer field that the object API left unset now
raises a
ValueErrornaming the field rather than failing later onNone. compute_tensor_ctyperejectsSTRINGtensors instead of returning a-1sentinel.- The test floor moved to
pytest>=9.1.1; pytest 8’sskip/failwrappers are not statically analyzable.
0.24.0 (2026-09-29)
Section titled “0.24.0 (2026-09-29)”Features
Section titled “Features”- add a knob registry and split the optimization plan and report (#521) (c7dba00), closes #517
- add the optimization section, the accumulation knob and a per-layer optimization plan (#516) (57a80e7), closes #513
- emit NT_N_PACKED weights for float 1x1 convolution and fully connected (#508) (e119e23), closes #507
Bug Fixes
Section titled “Bug Fixes”- bump the docs reference exports with each release (#505) (58c2aac), closes #504
- infer reduction shapes from explicit AIR metadata (#509) (d62223e), closes #479
- keep the docs exports stable under release-please’s version bump (#526) (1a0874f), closes #525
- restore neutral active top navigation (#514) (8d861db)
Documentation
Section titled “Documentation”- migrate site navigation, guides and reference content (#473) (f1a3ea6)
- polish AOT hero and product navigation (#518) (70e3edc)
0.23.0 (2026-09-26)
Section titled “0.23.0 (2026-09-26)”⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- generated modules, integer-only ones included, no longer build against ns-cmsis-nn v7.32.x to v7.34.x.
Features
Section titled “Features”- add greedy-by-size and hill-climb experimental memory planners (#359) (6ac4d16)
- declare the public Python API with all and an API manifest (d2661c3), closes #455
- dispatch float PACK, UNPACK and SPLIT to native ns-cmsis-nn kernels (#429) (20f407b)
- export the configuration, CLI, target, error and module-layout reference data (b252649)
- place operator code in ITCM and document ITCM tensor placement (#493) (849b4ab)
- resolve shape expressions in the shape propagation fixed point (#476) (e537d29)
- route 1D dilated depthwise conv to the optimized ns-cmsis-nn kernels (#503) (49a5ee4), closes #502
- support FP16 and FP32 nearest-neighbor resize (#404) (191bed4)
Bug Fixes
Section titled “Bug Fixes”- accept JSON values for structured list flags (#484) (0c5a6f0), closes #481
- accept scalar reduce axes and honor keepDims for integer extrema (#446) (fa0a562)
- carry SUM keepDims and require the exact output shape (#475) (6217d3e)
- drop the obsolete MAX/MIN/CLAMP undef prologue from the Zephyr test case (#478) (fd074a4), closes #305
- fail fast on out-of-range tensor dims in the buffer sizers (#477) (2db5429)
- infer dynamic shapes and fold static shape expressions (#439) (81f0076)
- keep emitted operator sources clean under -Wunused-parameter (#491) (ddeebcd), closes #406
- match attribute rule types case-insensitively (#482) (8077a0d), closes #474
- preserve INT16 hard swish prescale precision (#468) (ec917d1)
- preserve INT8 HARD_SWISH precision when deriving prescale (d68e851)
- raise on undefined template references during codegen (#423) (494aa30)
- remove the conversion work directory when the conversion ends (#483) (2a9be27), closes #467
- require ns-cmsis-nn v7.35.0 for every generated module (#498) (76e6961), closes #496
- scan registry discovery namespaces to a fixed point (#489) (9eaecda), closes #488
- tidy conversion edge cases around memory constraints, work dirs and golden data (#495) (ba3acc8), closes #485 #486 #487 #490
- warn when SVDF per-channel quantization is ignored (#494) (8502692), closes #329
Documentation
Section titled “Documentation”- consolidate strict template contract guidance (798a475)
- generate the Python, configuration, CLI, target and module reference (21e2068)
- scaffold the Astro site in docs/ and relocate MkDocs sources to mkdocs/ (6f55939), closes #454
0.22.0 (2026-09-15)
Section titled “0.22.0 (2026-09-15)”Features
Section titled “Features”- add explicit input and output byte-count macros (#426) (c0926c3)
- add native float argmin and argmax support (bf0fd47)
- add native float gather and extrema support (d5cb054)
- adopt native FP16 and FP32 RSQRT kernels (#425) (ce6c433)
- adopt native FP16 SQRT from CORE 7.33 (#421) (83c9f3d)
- lower non-scalar float SUB broadcasts to native kernels (0820177)
- support native float HARD_SWISH (#410) (20ea2ce)
Bug Fixes
Section titled “Bug Fixes”- declare FP16 utility type dependencies (#422) (958432e)
- expand per-tensor convolution quantization to every output channel (#411) (1375535)
- finish typed parser errors and quantization hints (8d65c66)
- honor precise float requirements in CMake and NSX (e546fc7)
- isolate generated layer parameters across model modules (#408) (6140aff)
- preserve integer mean logical output shapes (edd88be)
- reject incompatible integer fully connected filters (a37f767)
- reject int16 FULLY_CONNECTED filters no int16 kernel can honour (#418) (3c35331)
- reject invalid concatenation axes before code generation (9ea546a)
0.21.0 (2026-09-08)
Section titled “0.21.0 (2026-09-08)”Features
Section titled “Features”- lower native float MEAN and broadcast MUL (19d4d34)
- lower native float MEAN and broadcast MUL (19d4d34)
Bug Fixes
Section titled “Bug Fixes”0.20.0 (2026-09-07)
Section titled “0.20.0 (2026-09-07)”⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- generated modules require ns-cmsis-nn v7.32.0 or newer; the generated common header fails to compile against older releases.
Features
Section titled “Features”- accept float16/float32 in SQUEEZE, FILL, ZEROS_LIKE, and DILATE with an operator-named dtype error (#353) (3520a96)
- add float16 rolled and unrolled GRU support (4c3e974)
- e2e: report passing tolerance margins (#333) (259c763)
- wire float16/float32 SUB, STRIDED_SLICE, SLICE, and float16 SPLIT to the ns-cmsis-nn float kernels (#352) (d15ac41)
Bug Fixes
Section titled “Bug Fixes”- align generated code with ns-cmsis-nn 7.32.0 contracts (#386) (edc5176)
- avoid shared operator attribute defaults (f42a245)
- declare FP16 on every Cortex-M55 target so list-targets matches the supports_fp16 gate (#351) (bc2a2ea)
- e2e: migrate to CMSIS 6 and Cortex DFP (#332) (a10be71)
- keep config-derived custom platforms run-local instead of registering them globally (#334) (188660b)
- pin ns-cmsis-nn at v7.32.0 and size the float depthwise scratch like its sizer (#394) (e2e10f7)
- raise ConfigValueError with a hint for every user-facing operator validation failure (#400) (b2bb81e)
- reject float16 and float32 on the six operators ns-cmsis-nn ships integer kernels for (#396) (14227a8)
- report aot_tensor_io_t.size in bytes for non-int8 I/O tensors (#380) (9726d4e)
0.19.0 (2026-09-02)
Section titled “0.19.0 (2026-09-02)”⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- the
at110platform name is nowatomiq110;--platform.name at110no longer resolves.at110shipped in v0.18.0, but Atomiq is not in production and nothing downstream has pinned the name yet.
Features
Section titled “Features”- add atomiq110 board support and gate Ethos-U dispatch on NPU capability (#276) (8cee4b5)
- add broadcast_to op support (bc539f2)
- add broadcast_to op support (bc539f2)
- add dynamic_update_slice op support (13980ed)
- add dynamic_update_slice op support (13980ed)
- add mirror_pad op support (b46bb4b)
- add mirror_pad op support (b46bb4b)
- add RESIZE_BILINEAR and DILATE operators for int8/int16x8 (f052d51)
- add reverse_sequence op support (68e2d6b)
- add reverse_sequence op support (68e2d6b)
- add scatter_nd op support (70b9e40)
- add scatter_nd op support (70b9e40)
- add select_v2 op support (c9ad0eb)
- add select_v2 op support (c9ad0eb)
- add tile op support (1570971)
- add tile op support (1570971)
- add where op support (6ec8a75)
- add where op support (6ec8a75)
- aot: add float kernel support across operator pipeline (#246) (73d9138)
- carry recurrent state through the float LSTM path (a41df71)
- carry recurrent state through the unidirectional sequence LSTM path (f3777b6)
- carry recurrent state through the unidirectional sequence LSTM path (f3777b6)
- ci: file an issue when the weekly release-model E2E run fails (#299) (3fbe40b)
- cli: console presentation layer for conversion runs (#264) (1066679)
- cli: help panels, factory defaults, quick-start epilog, banner (#261) (c263d26)
- config: warn on unknown nested config keys with did-you-mean (#268) (6363014)
- declare apollo510l compatibility in nsx module manifest (6c4fb78)
- document and expose supported HeliaAOT target names (3b5880a)
- document and expose supported HeliaAOT target names (3b5880a)
- expose supported target names via CLI and add M55 targets (fc570ba)
- float: enable FP16/FP32 for ABS, PRELU, and SUM via ns-cmsis-nn float kernels (#279) (2be2bd5)
- gate releases on manifest-backed model-corpus E2E (#277) (0c41e19)
- guard resolve against NULL ctx.buf in kernels that dereference it (#325) (9abc3b0), closes #316
- rename Apollo330P target and match target names case-insensitively (3cb2215)
- support Python 3.13 and 3.14 (8644efa)
- support stateful unidirectional sequence LSTM (b0449a1)
- tensor attribute-key warnings and programmatic API surface (#270) (370d954)
- typed errors with hints and clean CLI exit codes (#263) (7f031fa)
Bug Fixes
Section titled “Bug Fixes”- Add missing doxygen blocks to header (8700eb0)
- add, mul, sub, expand_dims, max, min, gather bugs (dbd7a27)
- address Copilot review feedback on shape propagation (bd97015)
- adopt context-only run signature in dilate and resize_bilinear templates (421340c)
- align svdf recurrent state validation (#290) (bda8750)
- aot: narrow extern “C” guard scope in generated headers (#250) (a5154f1)
- block unresolved dims in shape rules and guard the memory planner (0d274ba)
- build the CLI on Python 3.14 (c29139c)
- Correct broadcast_to and batch_to_space rules with proper conditions (8a0b054)
- count rule-confirmed shapes as resolved and claim rank-4-only STB/BTS (3208948)
- e2e: clone the public ns-cmsis-nn without a credential (#291) (a22304e)
- e2e: seed Keras initializers so bidirectional-GRU weights are reproducible (#314) (116f176), closes #301
- emit scalar gather indices as rank-1 for heliaCore kernel contract (3399a8d)
- fix add, mul, sub, expand_dims, max, min, gather bugs (dbd7a27)
- fix add, mul, sub, expand_dims, max, min, gather bugs (4f41a11)
- Handle potential edge case where input batch shape is also negative (not just shape signature) (90d1021)
- harden rank-0 and expand_dims emission found by adversarial review (7611299)
- harden RESIZE_BILINEAR/DILATE validation, correct rounding, drop broken MVE path (95b4199)
- Improve Windows compatibility (d0cffaa)
- keep includes outside extern “C” guard in new header templates (1d7e2d4)
- make e2e stimulus deterministic (#288) (101aec9)
- make generated Zephyr test case CMSIS 6 compatible (#287) (0cc6e13)
- make shape propagation rule ownership explicit (722b642)
- make shape propagation rule ownership explicit (722b642)
- name custom-platform fields in both YAML and CLI spellings (160e7ab)
- normalize negative concat axis before the ownership test (e7cf8ff)
- raise the ns-cmsis-nn floor to 7.31.0 and pin e2e to match (#340) (59f4c6f), closes #304
- Reject index depth == 0 (114a85b)
- reject SVDF weights_feature/bias dtypes the kernels cannot consume (#312) (90c661e)
- release: publish the GitHub Release only after model-corpus qualification (#298) (85eda59)
- Remove _PUT_IN_MRAM_INIT to const mapping in test_emit-stage.py (a765e04)
- remove const from MRAM init placement macro (3e3d519)
- remove const from MRAM init placement macro (3e3d519)
- remove const from MRAM init placement macro (dafd5b4)
- remove duplicate numpy import in parser tests (a8dc4a7)
- replace in-place ndarray shape assignment (6edada0)
- resolve merge conflicts with main (e14c183)
- resolve merge conflicts with main (305d13f)
- restore CMSIS-NN no-clip sentinel and require persistent LSTM state (2641268)
- restore scalar-scalar support for maximum and minimum (0a54eb0)
- size int8 batch-matmul scratch from the RHS row count (#296) (28636e9)
- size transpose-conv int8 scratch so the kernel gets a real buffer (#313) (bfe683f)
- update fp16 golden test for the renamed LSTM e2e generator (9bdf93b)
- validate SPLIT output shapes against the split lengths before code generation (#344) (44bd27c), closes #322
- validate SUM output shape, rank, and constant axis against the kernel contract (#326) (b82b256), closes #280 #207
- validate SVDF shapes against the ns-cmsis-nn kernel contract (#324) (9e3535d), closes #317 #280
- write generated files as UTF-8 with LF newlines (4508272)
Documentation
Section titled “Documentation”- document exact-rational vs TFLite 10-bit divergence bounds (a6ab417)
- drop interpreter pin from pipx example (3c87204)
- ethos-u: correct driver contract and document known gaps (#266) (364cf21)
- tighten recurrent-state wording and pass byte length to arm_memset_s8 (4d30064)
- wire agent instructions into every tool and add operational guidance (#272) (c68882a)
Earlier releases are in CHANGELOG.md.