# Memory placement

Choose where model data lives and how the application supplies its storage. Start with the [memory overview](https://ambiqai.github.io/helia-aot/guide/memory/) if arena roles are new to you.

## Start with the default plan

First convert with your real target and enable a residency report. The default
planner and internally allocated arenas need no placement rules:

```yaml
memory:
  dump_residency_json: true
```

Read `used`, `memory` and `role` in `<prefix>_residency.json`. Add a budget or
placement rule only after identifying the bank or tensor you need to change.
Reconvert and check that tensor's report row; a successful conversion with an
unchanged row may mean the rule did not match. Rebuild and validate whenever
placement changes. The report describes the model's plan; the linker map must
also account for your application, stacks, heaps and other modules.

## Choosing where a tensor goes

Placement is set by tensor attribute rules under `memory.tensors`, using the
precedence described in
[Configuring a conversion](https://ambiqai.github.io/helia-aot/guide/configuring/#precedence).

| Memory | Typical use |
| --- | --- |
| `ITCM` | Code or explicitly placed tensors on a target that exposes ITCM. Share the budget with linked code; see [ITCM placement](https://ambiqai.github.io/helia-aot/guide/itcm-placement/). |
| `DTCM` | Scratch and activations when a tightly coupled data bank is available. |
| `SRAM` | General-purpose writable data within the application budget. |
| `PSRAM` | Large working sets, off chip. Volatile. |
| `MRAM` | Non-volatile constants. Read-only, so never scratch or persistent. |
| `DRAM` | External bulk storage, where the part has it. |

Which of these a target actually has is in the
[Targets reference](https://ambiqai.github.io/helia-aot/reference/targets/), and the platform decides
how a memory maps onto a linker section.

A rule's `type` is a tensor kind, and the three kinds are matched by their
lowercase spellings: `constant`, `persistent` and `scratch`. The comparison is
exact, so `type: CONSTANT` matches nothing and the rule is silently skipped.
For operator rules, use canonical uppercase names such as `CONV_2D`: the
operator entity is canonicalized before matching, but the rule's `type` string
is still compared exactly.

Start broad and refine. A global rule sets the default and a kind-scoped rule
moves a class of tensors. A kind-only rule outranks an id-only rule; use both
`type` and `id` for an exception within a kind.

```yaml
memory:
  tensors:
    - attributes: { memory: DTCM }          # default everything to DTCM
    - type: persistent
      attributes: { memory: SRAM }          # state does not need to be hot
    - type: constant
      attributes: { memory: MRAM }          # weights stay in non-volatile flash
    - type: scratch
      id: ["hot_buf_0", "hot_buf_1"]
      attributes: { memory: DTCM }          # two buffers that must stay fast
```

### A placed plan

The committed documentation fixture was converted under the rules below,
so the view shows a converter-produced plan. A rule's `type` is the tensor
kind as the residency report spells it,
and its `id` is the tensor id the same report carries; `11` is the selected constant tensor in this example, which is staged while the other constants stay cold; see
[Cold and staged constants](https://ambiqai.github.io/helia-aot/guide/memory-placement/#cold-and-staged-constants).

```yaml title="memory placement, from residency-example-placed.json"
memory:
  tensors:
    - attributes:
        memory: DTCM
    - type: constant
      attributes:
        memory: MRAM
    - id: "11"
      attributes:
        constant_destination_memory: SRAM
        memory: MRAM
```

**Placed.** The rules above spread the plan across 3 memories on `apollo510_evb`: `dtcm` holds the persistent arena and the scratch arena, `sram` holds the staged constant arena and `mram` holds the cold constant arena.

## Constraints

`memory.constraints` caps a memory or strengthens its alignment:

```yaml
memory:
  constraints:
    - name: DTCM
      max_size: 131072
      arena_alignment: 32
    - name: PSRAM
      arena_alignment: 64
```

`arena_alignment` applies both to the arena's base symbol and to every tensor
slot inside it, which is what DMA engines and cache-line-coherent transfers
need. It defaults to the larger of 16 bytes and the platform's own minimum
alignment.

Choose `max_size` from the budget your application can give this model, not
from the bank's total capacity alone. The planner accounts for its arena
roles within that budget; it cannot see allocations made elsewhere in your
firmware. If a pinned tensor cannot fit, reduce the working set, choose another
available bank or revise the budget. Compare the resulting report and linker
map before deploying.

For example, a neuralSPOT-X application on an Apollo510 EVB keeps its own
stack, heap and data in the same 496 KB TCM region as the module; an AD01 build
used about 46 KB of it. Leave that room by capping DTCM, and list the other
memories so the rest of the model can still spill into them:

```yaml
memory:
  constraints:
    - name: DTCM
      max_size: 458752   # 507904 B of MCU_TCM less 48 KB for the application
    - name: SRAM
    - name: MRAM
```

An explicit constraint list is the complete set of memories the planner may
use: on `apollo510_evb`, add `- name: PSRAM` to keep PSRAM available.

## Cold and staged constants

A constant is **cold** when its source and runtime memory are the same:
one read-only blob, read in place, with no hydration copy. Its memory cost is
charged to that bank. A cold constant placed in writable DTCM still occupies
DTCM; “cold” does not imply non-volatile storage or zero RAM use.

A constant is **staged** when `constant_destination_memory` names a different
memory. Then a contiguous source blob lives in the cold memory and a writable
runtime arena is filled from it before inference.

```yaml
memory:
  tensors:
    - type: constant
      attributes:
        memory: MRAM                      # where the bytes are stored
        constant_destination_memory: DTCM # where the kernels read them
```

Staging trades RAM for a writable copy of the weights and a one-time transfer,
in exchange for control over where the kernels read from. The source blob
also remains part of the module's storage footprint.

| Situation | Choose |
| --- | --- |
| Weights fit in non-volatile flash the core can execute and read in place | Cold. Simplest, and no RAM cost. |
| Weights sit in external storage with a high per-access cost, or no execute-in-place window | Staged, into SRAM or DTCM. |
| Repeated weight reads appear to dominate a measured layer | Try staging into DTCM or SRAM within budget, then validate and measure again. |
| Tiny shared biases and scalars | Start cold; stage only if the measured benefit justifies the additional copy. |

The effect depends on the kernel, memory system and working set. Compare
whole-inference time as well as the added runtime bytes and initialization cost.

### The hydration contract

When at least one constant is staged, the module emits a weak
`<prefix>_hydrate_constants()` whose default body is one `memcpy` per staged
arena. `<prefix>_model_init()` calls it between `<prefix>_context_init()` and
the operator init loop, so staged arenas are populated before any kernel can
read them. The normal sequence is:

1. Bind every arena if the application owns the buffers.
2. Call `<prefix>_model_init(&ctx)` and check its return value. It initializes
   the context, zeroes persistent storage, hydrates constants, then initializes
   operators. Only a successful initialization enables `model_run`.
3. Populate model inputs and call `<prefix>_model_run(&ctx)`, checking its
   return value on every invocation.

For DMA, decompression or preloaded data, replace the weak hydrate function
with a strong definition. The replacement must make all staged bytes ready
before returning success and call `<prefix>_mark_hydrated()`. Operator init
hooks may read weights immediately after that return. An asynchronous transfer
must therefore be completed or waited for inside this initialization sequence.

Every `model_init` calls `context_init`, which clears the hydration latch.
Calling the default helper, or marking the latch, *before* `model_init` does
not avoid the later copy. The helper is idempotent only until the next latch
reset. To reuse already staged bytes, make that decision in your replacement
helper at the fixed call point above.

A run before successful model initialization returns status `14`
(`not_initialized`). After successful initialization, a cleared hydration
latch returns `200` (`hydration_required`). Reinitialization resets persistent
state as well as the latch. `memory.auto_hydrate_constants` is deprecated;
it does not change this sequence.

### What the planner guarantees, and what it does not

Guaranteed, per arena:

- one source memory per destination arena, so the source blob is one
  contiguous symbol and hydration is one bulk transfer;
- matching relative offsets between source and destination;
- alignment padding charged to both sides, so the transfer neither undershoots
  nor overshoots.

Not supported: two destination arenas backed by the same memory, reusing a
scratch arena to hold constants between inferences, and lazy or on-demand
hydration. Hydration is one shot.

## Caller-supplied arenas

With `memory.allocate_arenas: false` the module declares the arenas but
allocates none of them. The application owns every buffer and binds each one
before `<prefix>_model_init()`.

The header gives you a size macro, an alignment macro and a region enumerator
per arena, plus `<prefix>_bind_arena()` and `<prefix>_bind_arenas()`. Binding
rejects an unknown region, a null pointer, an undersized buffer and a
misaligned one, each with its own status code; the signatures and the codes
are in the
[Generated module reference](https://ambiqai.github.io/helia-aot/reference/module/#with-external-arenas).

The bind table and hydration latch are module-global. Two context structs for
the same generated prefix do not create two independent model instances.
Use separately prefixed modules for independent state; sharing their scratch
storage is possible only when their live data does not overlap in time.

In this mode you must bind cold constant arenas too, to the bytes loaded from
the emitted constant blob sidecar. The built-in test case is not a sidecar
loader; use a [sidecar-aware test harness](https://ambiqai.github.io/helia-aot/guide/testing/#external-arena-tests)
when validating that configuration.

### Two models in one image

Every generated symbol carries `module.prefix`, so two modules converted with
different prefixes coexist in one binary with two independent bind tables. A
firmware image with a wake-word model and a classifier, both converted with
`memory.allocate_arenas: false`, is bound like this.

1. Convert each model with its own prefix, for example `--module.prefix wake`
   and `--module.prefix clf`. Two modules that share a prefix collide at link
   time.
2. Allocate the buffers. Sizes and alignments come from the generated headers:
   `wake_arena_sizes[]` and `wake_arena_alignments[]`, indexed by the
   `wake_arena_region_t` enumerators, and the same pair under `clf_`.
3. Bind every region of each module before that module is initialized, with
   `wake_bind_arena(region, buffer, size)` per region or `wake_bind_arenas()`
   in one call. Binding rejects an unknown region, a null pointer, an
   undersized buffer and a misaligned one with a distinct status each. In this
   mode the cold constant regions are bound too, to the bytes of the emitted
   `wake_arena_const_<memory>__blob.bin` sidecars.
4. Call `wake_model_init(&wake_ctx)` and `clf_model_init(&clf_ctx)`. Each one
   validates its own bind table, zeroes its persistent arena, hydrates any
   staged constant arena, then runs the per-operator init hooks. An unbound
   region fails here rather than at the first inference.
5. Run either context with `wake_model_run` or `clf_model_run`. A run touches
   only the arenas bound for that prefix.

The persistent arenas have to stay separate, because they hold each model's
state for the life of the program. The scratch arenas are the ones worth
sharing: one buffer, sized to the larger of `wake_arena_<memory>_size` and
`clf_arena_<memory>_size` and aligned to the larger of the two alignments, can
be bound into both.

:::caution
Model inputs and outputs are ordinary planned tensors, so they live in the
scratch arena unless a rule moves them. Two models sharing a scratch buffer
must not have live data in it at the same time: copy the first model's outputs
somewhere else before you run the second, or give the two models their own
scratch. A tensor rule moves a tensor between memories; it cannot change a
tensor's kind, so there is no rule that makes an output outlive the run.
:::

## Next

[Inspect the residency report](https://ambiqai.github.io/helia-aot/guide/memory-reports/) to check the emitted plan.
