Memory
Memory planning assigns tensor storage before firmware is built. Tensor lifetimes let scratch storage be reused; placement rules choose memory banks and enforce your budgets.
Memory planning
Section titled “Memory planning”- Prune and intern: remove unread tensors and share identical constants.
- Memory planner: turn lifetimes, rules and constraints into offsets.
- Arena roles: keep scratch, persistent and constant storage distinct.
- Render plan: pass the layout and its hashes to code generation.
Arenas, not symbols
Section titled “Arenas, not symbols”Every tensor, scratch, persistent and constant alike, lives in a per-memory arena. There are no per-tensor static symbols. Each tensor descriptor carries a region, an offset and a size, resolved at run time against the arena buffer for that region.
| Role | What it holds | Lifetime and initialization |
|---|---|---|
scratch |
Activations and operator working buffers | Slots are reused across operators by liveness. Not initialized. |
persistent |
Variable tensors and resource variables | Whole program. Zeroed explicitly in context_init. |
constant |
Weights and other read-only data | Whole program. Either read in place, or hydrated into a writable arena. |
Choose the next control
Section titled “Choose the next control”Start with a successful conversion for your actual target and enable
memory.dump_residency_json. The report tells you which bank each tensor uses
and how much of each arena is occupied. A generated module can fit its model
budget while the whole application still exceeds a linker region, so inspect
both the report and the firmware’s linker map.
| Need | Next step | Check the result |
|---|---|---|
| Understand the current footprint | Read the residency report | Inspect occupied bytes, bank names and both layout hashes. |
| Reserve room for the application | Set a memory budget | The conversion fits the limit; the complete firmware also fits its linker regions. |
| Put a tensor in another bank | Apply placement rules | Its report row shows the intended memory; numerical validation still passes. |
| Reduce scratch footprint | Compare planners | Compare occupied bytes and bank changes under identical budgets. |
| Copy weights to a runtime bank | Stage constants | Account for source and destination storage, initialization cost and measured inference time. |
| Supply or share application buffers | Bind external arenas | Every region is bound before initialization; shared scratch has no overlapping live data. |
Ownership and initialization
Section titled “Ownership and initialization”By default the generated module allocates its arena buffers. Setting
memory.allocate_arenas: false transfers buffer ownership to the application,
which must bind every region before model_init. Placement rules move tensors
between banks; they do not change a tensor’s role or extend an output’s lifetime.
Inputs and outputs can occupy reusable scratch storage.
model_init initializes the context, resets persistent state, hydrates staged
constants and runs operator initialization. Check its result before calling
model_run. Reinitializing resets state; it is not the way to begin every
iteration of a streaming inference sequence.
A context struct refers to the module’s arena storage. Allocating another context struct for the same generated prefix does not allocate independent buffers or an independent hydration latch. For two separate models, use unique prefixes and follow the buffer-sharing rules.