Skip to content
heliaAOT
HELIA HUB

Memory

Memory planning assigns tensor storage before firmware is built. Tensor lifetimes let scratch storage be reused; placement rules choose memory banks and enforce your budgets.

  1. Prune and intern: remove unread tensors and share identical constants.
  2. Memory planner: turn lifetimes, rules and constraints into offsets.
  3. Arena roles: keep scratch, persistent and constant storage distinct.
  4. Render plan: pass the layout and its hashes to code generation.

Every tensor, scratch, persistent and constant alike, lives in a per-memory arena. There are no per-tensor static symbols. Each tensor descriptor carries a region, an offset and a size, resolved at run time against the arena buffer for that region.

Role What it holds Lifetime and initialization
scratch Activations and operator working buffers Slots are reused across operators by liveness. Not initialized.
persistent Variable tensors and resource variables Whole program. Zeroed explicitly in context_init.
constant Weights and other read-only data Whole program. Either read in place, or hydrated into a writable arena.

Start with a successful conversion for your actual target and enable memory.dump_residency_json. The report tells you which bank each tensor uses and how much of each arena is occupied. A generated module can fit its model budget while the whole application still exceeds a linker region, so inspect both the report and the firmware’s linker map.

Need Next step Check the result
Understand the current footprint Read the residency report Inspect occupied bytes, bank names and both layout hashes.
Reserve room for the application Set a memory budget The conversion fits the limit; the complete firmware also fits its linker regions.
Put a tensor in another bank Apply placement rules Its report row shows the intended memory; numerical validation still passes.
Reduce scratch footprint Compare planners Compare occupied bytes and bank changes under identical budgets.
Copy weights to a runtime bank Stage constants Account for source and destination storage, initialization cost and measured inference time.
Supply or share application buffers Bind external arenas Every region is bound before initialization; shared scratch has no overlapping live data.

By default the generated module allocates its arena buffers. Setting memory.allocate_arenas: false transfers buffer ownership to the application, which must bind every region before model_init. Placement rules move tensors between banks; they do not change a tensor’s role or extend an output’s lifetime. Inputs and outputs can occupy reusable scratch storage.

model_init initializes the context, resets persistent state, hydrates staged constants and runs operator initialization. Check its result before calling model_run. Reinitializing resets state; it is not the way to begin every iteration of a streaming inference sequence.

A context struct refers to the module’s arena storage. Allocating another context struct for the same generated prefix does not allocate independent buffers or an independent hydration latch. For two separate models, use unique prefixes and follow the buffer-sharing rules.