Memory planners
Start with the default greedy planner. An alternative is useful when the
residency report shows a scratch footprint you need to reduce. Keep the model,
platform, placement rules and budgets fixed while comparing planners; a lower
footprint alone does not establish faster inference.
Compare a candidate
Section titled “Compare a candidate”Add this block to a configuration that already converts:
memory: planner: hill_climb planner_options: iterations: 300 seed: 0 max_stall_iterations: null dump_residency_json: trueConvert to a separate output directory for each planner. Compare scratch
used, peak_live and fragmentation in the
residency reports, then compare the memory
bank of each tensor. A successful candidate fits the same constraints and
passes the same numerical validation as the baseline. Measure on your target
before choosing a layout for latency.
These are the hill-climb defaults: at most 300 order perturbations, deterministic
seed 0 and no early stall limit. iterations: 0 uses its greedy-by-size seed
layout. A positive max_stall_iterations stops after that many consecutive
attempts without improvement. More iterations cost conversion time; they do
not increase the model’s runtime search work because there is no runtime
planning.
The planner
Section titled “The planner”memory.planner names the algorithm that turns tensor lifetimes into arena
offsets. Three are registered; one is the default and two are experimental,
so an existing conversion keeps its layout until you opt into one of them.
| Value | What it does |
|---|---|
greedy |
The default. First fit in lifetime-start order over a shared buffer per writable memory: a tensor takes the first gap large enough, and gaps left by tensors whose lifetime has ended are reused. |
greedy_by_size |
Experimental. Packs offsets largest-first, which usually lowers the scratch high-water mark on models whose tensors differ a lot in size. |
hill_climb |
Experimental. A local search over greedy-by-size placement orders; its total scratch high-water mark across banks is never worse than its greedy_by_size seed. Tuned through memory.planner_options (iterations, seed, max_stall_iterations). |
memory.planner_options is a JSON object validated against the selected
planner’s option schema, so a key the planner does not know is a
configuration error, not a silent no-op. Every planner honours the same
placement rules: per-tensor memory pins, constant_destination_memory
routing and per-memory max_size budgets. The outcomes can still differ,
because unpinned scratch spills to the next writable memory when the current
one fills, and a different visit order can move a tensor between memories,
which can change access latency as well as footprint. Hill-climb minimizes
the total high-water mark, breaking ties by per-bank usage in the preferred
memory order. This does not guarantee that every bank uses fewer bytes.
Its greedy-by-size seed must fit first: an infeasible seed fails the conversion
before the search starts.
The greedy planner walks the operators in the order the model gives them, so a conversion is reproducible: the same model, target and rules produce the same offsets and the same plan hash. It fills the three arena families independently. Scratch arenas are reused by liveness, persistent arenas are bump allocated and never reclaimed, and constant arenas are keyed by the memory the kernels read them from.
Alignment is resolved per tensor as the largest of three floors: the platform’s
min_alignment, the natural alignment of the element type, and any stronger
hint the operator asked for, which is how dot-product kernels and Ethos-U
command streams get the boundaries they need.
Planners are a registered extension point rather than a fixed list. A planner
is a subclass of AirMemoryPlanner with a NAME, registered through
register_memory_planner, and the emitter reaches it through the structural
protocol in helia_aot.memory.backend, so a planner published in another
package does not have to import the code generator to be usable. The published memory API reference covers the planning data contracts; the planner implementation and registration internals are not part of that public API manifest.