# Memory

Memory planning assigns tensor storage before firmware is built. Tensor lifetimes let scratch storage be reused; placement rules choose memory banks and enforce your budgets.

- [Placement and application buffers](https://ambiqai.github.io/helia-aot/guide/memory-placement/): Choose memory banks, stage constants and bind buffers owned by your application.
- [Planner options](https://ambiqai.github.io/helia-aot/guide/memory-planners/): Use the default planner or evaluate the two experimental alternatives.
- [Residency reports](https://ambiqai.github.io/helia-aot/guide/memory-reports/): Inspect arena usage, tensor offsets and layout hashes.

## Memory planning

1. **Prune and intern**: remove unread tensors and share identical constants.
2. **Memory planner**: turn lifetimes, rules and constraints into offsets.
3. **Arena roles**: keep scratch, persistent and constant storage distinct.
4. **Render plan**: pass the layout and its hashes to code generation.

## Arenas, not symbols

Every tensor, scratch, persistent and constant alike, lives in a per-memory
**arena**. There are no per-tensor static symbols. Each tensor descriptor
carries a region, an offset and a size, resolved at run time against the
arena buffer for that region.

| Role | What it holds | Lifetime and initialization |
| --- | --- | --- |
| `scratch` | Activations and operator working buffers | Slots are reused across operators by liveness. Not initialized. |
| `persistent` | Variable tensors and resource variables | Whole program. Zeroed explicitly in `context_init`. |
| `constant` | Weights and other read-only data | Whole program. Either read in place, or hydrated into a writable arena. |

:::note
`context_init` writes raw zero bytes. For an asymmetric int8 tensor, real
zero is the zero point, not the byte 0, so operators whose state is read back
as a live input re-seed it in their init hook. Operators whose kernel cannot
represent a non-zero state zero point reject it instead.
:::

## Choose the next control

Start with a successful conversion for your actual target and enable
`memory.dump_residency_json`. The report tells you which bank each tensor uses
and how much of each arena is occupied. A generated module can fit its model
budget while the whole application still exceeds a linker region, so inspect
both the report and the firmware's linker map.

| Need | Next step | Check the result |
| --- | --- | --- |
| Understand the current footprint | [Read the residency report](https://ambiqai.github.io/helia-aot/guide/memory-reports/) | Inspect occupied bytes, bank names and both layout hashes. |
| Reserve room for the application | [Set a memory budget](https://ambiqai.github.io/helia-aot/guide/memory-placement/#constraints) | The conversion fits the limit; the complete firmware also fits its linker regions. |
| Put a tensor in another bank | [Apply placement rules](https://ambiqai.github.io/helia-aot/guide/memory-placement/#choosing-where-a-tensor-goes) | Its report row shows the intended memory; numerical validation still passes. |
| Reduce scratch footprint | [Compare planners](https://ambiqai.github.io/helia-aot/guide/memory-planners/) | Compare occupied bytes and bank changes under identical budgets. |
| Copy weights to a runtime bank | [Stage constants](https://ambiqai.github.io/helia-aot/guide/memory-placement/#cold-and-staged-constants) | Account for source and destination storage, initialization cost and measured inference time. |
| Supply or share application buffers | [Bind external arenas](https://ambiqai.github.io/helia-aot/guide/memory-placement/#caller-supplied-arenas) | Every region is bound before initialization; shared scratch has no overlapping live data. |

## Ownership and initialization

By default the generated module allocates its arena buffers. Setting
`memory.allocate_arenas: false` transfers buffer ownership to the application,
which must bind every region before `model_init`. Placement rules move tensors
between banks; they do not change a tensor's role or extend an output's lifetime.
Inputs and outputs can occupy reusable scratch storage.

`model_init` initializes the context, resets persistent state, hydrates staged
constants and runs operator initialization. Check its result before calling
`model_run`. Reinitializing resets state; it is not the way to begin every
iteration of a streaming inference sequence.

A context struct refers to the module's arena storage. Allocating another
context struct for the same generated prefix does not allocate independent
buffers or an independent hydration latch. For two separate models, use unique
prefixes and follow the [buffer-sharing rules](https://ambiqai.github.io/helia-aot/guide/memory-placement/#two-models-in-one-image).
