Skip to content
heliaAOT
HELIA HUB

Memory placement

Choose where model data lives and how the application supplies its storage. Start with the memory overview if arena roles are new to you.

First convert with your real target and enable a residency report. The default planner and internally allocated arenas need no placement rules:

memory:
dump_residency_json: true

Read used, memory and role in <prefix>_residency.json. Add a budget or placement rule only after identifying the bank or tensor you need to change. Reconvert and check that tensor’s report row; a successful conversion with an unchanged row may mean the rule did not match. Rebuild and validate whenever placement changes. The report describes the model’s plan; the linker map must also account for your application, stacks, heaps and other modules.

Placement is set by tensor attribute rules under memory.tensors, using the precedence described in Configuring a conversion.

Memory Typical use
ITCM Code or explicitly placed tensors on a target that exposes ITCM. Share the budget with linked code; see ITCM placement.
DTCM Scratch and activations when a tightly coupled data bank is available.
SRAM General-purpose writable data within the application budget.
PSRAM Large working sets, off chip. Volatile.
MRAM Non-volatile constants. Read-only, so never scratch or persistent.
DRAM External bulk storage, where the part has it.

Which of these a target actually has is in the Targets reference, and the platform decides how a memory maps onto a linker section.

A rule’s type is a tensor kind, and the three kinds are matched by their lowercase spellings: constant, persistent and scratch. The comparison is exact, so type: CONSTANT matches nothing and the rule is silently skipped. For operator rules, use canonical uppercase names such as CONV_2D: the operator entity is canonicalized before matching, but the rule’s type string is still compared exactly.

Start broad and refine. A global rule sets the default and a kind-scoped rule moves a class of tensors. A kind-only rule outranks an id-only rule; use both type and id for an exception within a kind.

memory:
tensors:
- attributes: { memory: DTCM } # default everything to DTCM
- type: persistent
attributes: { memory: SRAM } # state does not need to be hot
- type: constant
attributes: { memory: MRAM } # weights stay in non-volatile flash
- type: scratch
id: ["hot_buf_0", "hot_buf_1"]
attributes: { memory: DTCM } # two buffers that must stay fast

The committed documentation fixture was converted under the rules below, so the view shows a converter-produced plan. A rule’s type is the tensor kind as the residency report spells it, and its id is the tensor id the same report carries; 11 is the selected constant tensor in this example, which is staged while the other constants stay cold; see Cold and staged constants.

memory placement, from residency-example-placed.json
memory:
tensors:
- attributes:
memory: DTCM
- type: constant
attributes:
memory: MRAM
- id: "11"
attributes:
constant_destination_memory: SRAM
memory: MRAM
DTCM
persistent256 bytes
scratch507,648 bytes
SRAM
constant, staged32,768 bytes
MRAM
constant, cold11,336 bytes
The residency report for the demo fixture model converted for apollo510_evb under the rules above.

Placed. The rules above spread the plan across 3 memories on apollo510_evb: dtcm holds the persistent arena and the scratch arena, sram holds the staged constant arena and mram holds the cold constant arena.

memory.constraints caps a memory or strengthens its alignment:

memory:
constraints:
- name: DTCM
max_size: 131072
arena_alignment: 32
- name: PSRAM
arena_alignment: 64

arena_alignment applies both to the arena’s base symbol and to every tensor slot inside it, which is what DMA engines and cache-line-coherent transfers need. It defaults to the larger of 16 bytes and the platform’s own minimum alignment.

Choose max_size from the budget your application can give this model, not from the bank’s total capacity alone. The planner accounts for its arena roles within that budget; it cannot see allocations made elsewhere in your firmware. If a pinned tensor cannot fit, reduce the working set, choose another available bank or revise the budget. Compare the resulting report and linker map before deploying.

For example, a neuralSPOT-X application on an Apollo510 EVB keeps its own stack, heap and data in the same 496 KB TCM region as the module; an AD01 build used about 46 KB of it. Leave that room by capping DTCM, and list the other memories so the rest of the model can still spill into them:

memory:
constraints:
- name: DTCM
max_size: 458752 # 507904 B of MCU_TCM less 48 KB for the application
- name: SRAM
- name: MRAM

An explicit constraint list is the complete set of memories the planner may use: on apollo510_evb, add - name: PSRAM to keep PSRAM available.

A constant is cold when its source and runtime memory are the same: one read-only blob, read in place, with no hydration copy. Its memory cost is charged to that bank. A cold constant placed in writable DTCM still occupies DTCM; “cold” does not imply non-volatile storage or zero RAM use.

A constant is staged when constant_destination_memory names a different memory. Then a contiguous source blob lives in the cold memory and a writable runtime arena is filled from it before inference.

memory:
tensors:
- type: constant
attributes:
memory: MRAM # where the bytes are stored
constant_destination_memory: DTCM # where the kernels read them

Staging trades RAM for a writable copy of the weights and a one-time transfer, in exchange for control over where the kernels read from. The source blob also remains part of the module’s storage footprint.

Situation Choose
Weights fit in non-volatile flash the core can execute and read in place Cold. Simplest, and no RAM cost.
Weights sit in external storage with a high per-access cost, or no execute-in-place window Staged, into SRAM or DTCM.
Repeated weight reads appear to dominate a measured layer Try staging into DTCM or SRAM within budget, then validate and measure again.
Tiny shared biases and scalars Start cold; stage only if the measured benefit justifies the additional copy.

The effect depends on the kernel, memory system and working set. Compare whole-inference time as well as the added runtime bytes and initialization cost.

When at least one constant is staged, the module emits a weak <prefix>_hydrate_constants() whose default body is one memcpy per staged arena. <prefix>_model_init() calls it between <prefix>_context_init() and the operator init loop, so staged arenas are populated before any kernel can read them. The normal sequence is:

  1. Bind every arena if the application owns the buffers.
  2. Call <prefix>_model_init(&ctx) and check its return value. It initializes the context, zeroes persistent storage, hydrates constants, then initializes operators. Only a successful initialization enables model_run.
  3. Populate model inputs and call <prefix>_model_run(&ctx), checking its return value on every invocation.

For DMA, decompression or preloaded data, replace the weak hydrate function with a strong definition. The replacement must make all staged bytes ready before returning success and call <prefix>_mark_hydrated(). Operator init hooks may read weights immediately after that return. An asynchronous transfer must therefore be completed or waited for inside this initialization sequence.

Every model_init calls context_init, which clears the hydration latch. Calling the default helper, or marking the latch, before model_init does not avoid the later copy. The helper is idempotent only until the next latch reset. To reuse already staged bytes, make that decision in your replacement helper at the fixed call point above.

A run before successful model initialization returns status 14 (not_initialized). After successful initialization, a cleared hydration latch returns 200 (hydration_required). Reinitialization resets persistent state as well as the latch. memory.auto_hydrate_constants is deprecated; it does not change this sequence.

What the planner guarantees, and what it does not

Section titled “What the planner guarantees, and what it does not”

Guaranteed, per arena:

  • one source memory per destination arena, so the source blob is one contiguous symbol and hydration is one bulk transfer;
  • matching relative offsets between source and destination;
  • alignment padding charged to both sides, so the transfer neither undershoots nor overshoots.

Not supported: two destination arenas backed by the same memory, reusing a scratch arena to hold constants between inferences, and lazy or on-demand hydration. Hydration is one shot.

With memory.allocate_arenas: false the module declares the arenas but allocates none of them. The application owns every buffer and binds each one before <prefix>_model_init().

The header gives you a size macro, an alignment macro and a region enumerator per arena, plus <prefix>_bind_arena() and <prefix>_bind_arenas(). Binding rejects an unknown region, a null pointer, an undersized buffer and a misaligned one, each with its own status code; the signatures and the codes are in the Generated module reference.

The bind table and hydration latch are module-global. Two context structs for the same generated prefix do not create two independent model instances. Use separately prefixed modules for independent state; sharing their scratch storage is possible only when their live data does not overlap in time.

In this mode you must bind cold constant arenas too, to the bytes loaded from the emitted constant blob sidecar. The built-in test case is not a sidecar loader; use a sidecar-aware test harness when validating that configuration.

Every generated symbol carries module.prefix, so two modules converted with different prefixes coexist in one binary with two independent bind tables. A firmware image with a wake-word model and a classifier, both converted with memory.allocate_arenas: false, is bound like this.

  1. Convert each model with its own prefix, for example --module.prefix wake and --module.prefix clf. Two modules that share a prefix collide at link time.
  2. Allocate the buffers. Sizes and alignments come from the generated headers: wake_arena_sizes[] and wake_arena_alignments[], indexed by the wake_arena_region_t enumerators, and the same pair under clf_.
  3. Bind every region of each module before that module is initialized, with wake_bind_arena(region, buffer, size) per region or wake_bind_arenas() in one call. Binding rejects an unknown region, a null pointer, an undersized buffer and a misaligned one with a distinct status each. In this mode the cold constant regions are bound too, to the bytes of the emitted wake_arena_const_<memory>__blob.bin sidecars.
  4. Call wake_model_init(&wake_ctx) and clf_model_init(&clf_ctx). Each one validates its own bind table, zeroes its persistent arena, hydrates any staged constant arena, then runs the per-operator init hooks. An unbound region fails here rather than at the first inference.
  5. Run either context with wake_model_run or clf_model_run. A run touches only the arenas bound for that prefix.

The persistent arenas have to stay separate, because they hold each model’s state for the life of the program. The scratch arenas are the ones worth sharing: one buffer, sized to the larger of wake_arena_<memory>_size and clf_arena_<memory>_size and aligned to the larger of the two alignments, can be bound into both.

Inspect the residency report to check the emitted plan.