Memory placement
Choose where model data lives and how the application supplies its storage. Start with the memory overview if arena roles are new to you.
Start with the default plan
Section titled “Start with the default plan”First convert with your real target and enable a residency report. The default planner and internally allocated arenas need no placement rules:
memory: dump_residency_json: trueRead used, memory and role in <prefix>_residency.json. Add a budget or
placement rule only after identifying the bank or tensor you need to change.
Reconvert and check that tensor’s report row; a successful conversion with an
unchanged row may mean the rule did not match. Rebuild and validate whenever
placement changes. The report describes the model’s plan; the linker map must
also account for your application, stacks, heaps and other modules.
Choosing where a tensor goes
Section titled “Choosing where a tensor goes”Placement is set by tensor attribute rules under memory.tensors, using the
precedence described in
Configuring a conversion.
| Memory | Typical use |
|---|---|
ITCM |
Code or explicitly placed tensors on a target that exposes ITCM. Share the budget with linked code; see ITCM placement. |
DTCM |
Scratch and activations when a tightly coupled data bank is available. |
SRAM |
General-purpose writable data within the application budget. |
PSRAM |
Large working sets, off chip. Volatile. |
MRAM |
Non-volatile constants. Read-only, so never scratch or persistent. |
DRAM |
External bulk storage, where the part has it. |
Which of these a target actually has is in the Targets reference, and the platform decides how a memory maps onto a linker section.
A rule’s type is a tensor kind, and the three kinds are matched by their
lowercase spellings: constant, persistent and scratch. The comparison is
exact, so type: CONSTANT matches nothing and the rule is silently skipped.
For operator rules, use canonical uppercase names such as CONV_2D: the
operator entity is canonicalized before matching, but the rule’s type string
is still compared exactly.
Start broad and refine. A global rule sets the default and a kind-scoped rule
moves a class of tensors. A kind-only rule outranks an id-only rule; use both
type and id for an exception within a kind.
memory: tensors: - attributes: { memory: DTCM } # default everything to DTCM - type: persistent attributes: { memory: SRAM } # state does not need to be hot - type: constant attributes: { memory: MRAM } # weights stay in non-volatile flash - type: scratch id: ["hot_buf_0", "hot_buf_1"] attributes: { memory: DTCM } # two buffers that must stay fastA placed plan
Section titled “A placed plan”The committed documentation fixture was converted under the rules below,
so the view shows a converter-produced plan. A rule’s type is the tensor
kind as the residency report spells it,
and its id is the tensor id the same report carries; 11 is the selected constant tensor in this example, which is staged while the other constants stay cold; see
Cold and staged constants.
memory: tensors: - attributes: memory: DTCM - type: constant attributes: memory: MRAM - id: "11" attributes: constant_destination_memory: SRAM memory: MRAMdemo fixture
model converted for apollo510_evb under the rules
above.
Placed. The rules above spread the plan across 3 memories on apollo510_evb: dtcm holds the persistent arena and the scratch arena, sram holds the staged constant arena and mram holds the cold constant arena.
Constraints
Section titled “Constraints”memory.constraints caps a memory or strengthens its alignment:
memory: constraints: - name: DTCM max_size: 131072 arena_alignment: 32 - name: PSRAM arena_alignment: 64arena_alignment applies both to the arena’s base symbol and to every tensor
slot inside it, which is what DMA engines and cache-line-coherent transfers
need. It defaults to the larger of 16 bytes and the platform’s own minimum
alignment.
Choose max_size from the budget your application can give this model, not
from the bank’s total capacity alone. The planner accounts for its arena
roles within that budget; it cannot see allocations made elsewhere in your
firmware. If a pinned tensor cannot fit, reduce the working set, choose another
available bank or revise the budget. Compare the resulting report and linker
map before deploying.
For example, a neuralSPOT-X application on an Apollo510 EVB keeps its own stack, heap and data in the same 496 KB TCM region as the module; an AD01 build used about 46 KB of it. Leave that room by capping DTCM, and list the other memories so the rest of the model can still spill into them:
memory: constraints: - name: DTCM max_size: 458752 # 507904 B of MCU_TCM less 48 KB for the application - name: SRAM - name: MRAMAn explicit constraint list is the complete set of memories the planner may
use: on apollo510_evb, add - name: PSRAM to keep PSRAM available.
Cold and staged constants
Section titled “Cold and staged constants”A constant is cold when its source and runtime memory are the same: one read-only blob, read in place, with no hydration copy. Its memory cost is charged to that bank. A cold constant placed in writable DTCM still occupies DTCM; “cold” does not imply non-volatile storage or zero RAM use.
A constant is staged when constant_destination_memory names a different
memory. Then a contiguous source blob lives in the cold memory and a writable
runtime arena is filled from it before inference.
memory: tensors: - type: constant attributes: memory: MRAM # where the bytes are stored constant_destination_memory: DTCM # where the kernels read themStaging trades RAM for a writable copy of the weights and a one-time transfer, in exchange for control over where the kernels read from. The source blob also remains part of the module’s storage footprint.
| Situation | Choose |
|---|---|
| Weights fit in non-volatile flash the core can execute and read in place | Cold. Simplest, and no RAM cost. |
| Weights sit in external storage with a high per-access cost, or no execute-in-place window | Staged, into SRAM or DTCM. |
| Repeated weight reads appear to dominate a measured layer | Try staging into DTCM or SRAM within budget, then validate and measure again. |
| Tiny shared biases and scalars | Start cold; stage only if the measured benefit justifies the additional copy. |
The effect depends on the kernel, memory system and working set. Compare whole-inference time as well as the added runtime bytes and initialization cost.
The hydration contract
Section titled “The hydration contract”When at least one constant is staged, the module emits a weak
<prefix>_hydrate_constants() whose default body is one memcpy per staged
arena. <prefix>_model_init() calls it between <prefix>_context_init() and
the operator init loop, so staged arenas are populated before any kernel can
read them. The normal sequence is:
- Bind every arena if the application owns the buffers.
- Call
<prefix>_model_init(&ctx)and check its return value. It initializes the context, zeroes persistent storage, hydrates constants, then initializes operators. Only a successful initialization enablesmodel_run. - Populate model inputs and call
<prefix>_model_run(&ctx), checking its return value on every invocation.
For DMA, decompression or preloaded data, replace the weak hydrate function
with a strong definition. The replacement must make all staged bytes ready
before returning success and call <prefix>_mark_hydrated(). Operator init
hooks may read weights immediately after that return. An asynchronous transfer
must therefore be completed or waited for inside this initialization sequence.
Every model_init calls context_init, which clears the hydration latch.
Calling the default helper, or marking the latch, before model_init does
not avoid the later copy. The helper is idempotent only until the next latch
reset. To reuse already staged bytes, make that decision in your replacement
helper at the fixed call point above.
A run before successful model initialization returns status 14
(not_initialized). After successful initialization, a cleared hydration
latch returns 200 (hydration_required). Reinitialization resets persistent
state as well as the latch. memory.auto_hydrate_constants is deprecated;
it does not change this sequence.
What the planner guarantees, and what it does not
Section titled “What the planner guarantees, and what it does not”Guaranteed, per arena:
- one source memory per destination arena, so the source blob is one contiguous symbol and hydration is one bulk transfer;
- matching relative offsets between source and destination;
- alignment padding charged to both sides, so the transfer neither undershoots nor overshoots.
Not supported: two destination arenas backed by the same memory, reusing a scratch arena to hold constants between inferences, and lazy or on-demand hydration. Hydration is one shot.
Caller-supplied arenas
Section titled “Caller-supplied arenas”With memory.allocate_arenas: false the module declares the arenas but
allocates none of them. The application owns every buffer and binds each one
before <prefix>_model_init().
The header gives you a size macro, an alignment macro and a region enumerator
per arena, plus <prefix>_bind_arena() and <prefix>_bind_arenas(). Binding
rejects an unknown region, a null pointer, an undersized buffer and a
misaligned one, each with its own status code; the signatures and the codes
are in the
Generated module reference.
The bind table and hydration latch are module-global. Two context structs for the same generated prefix do not create two independent model instances. Use separately prefixed modules for independent state; sharing their scratch storage is possible only when their live data does not overlap in time.
In this mode you must bind cold constant arenas too, to the bytes loaded from the emitted constant blob sidecar. The built-in test case is not a sidecar loader; use a sidecar-aware test harness when validating that configuration.
Two models in one image
Section titled “Two models in one image”Every generated symbol carries module.prefix, so two modules converted with
different prefixes coexist in one binary with two independent bind tables. A
firmware image with a wake-word model and a classifier, both converted with
memory.allocate_arenas: false, is bound like this.
- Convert each model with its own prefix, for example
--module.prefix wakeand--module.prefix clf. Two modules that share a prefix collide at link time. - Allocate the buffers. Sizes and alignments come from the generated headers:
wake_arena_sizes[]andwake_arena_alignments[], indexed by thewake_arena_region_tenumerators, and the same pair underclf_. - Bind every region of each module before that module is initialized, with
wake_bind_arena(region, buffer, size)per region orwake_bind_arenas()in one call. Binding rejects an unknown region, a null pointer, an undersized buffer and a misaligned one with a distinct status each. In this mode the cold constant regions are bound too, to the bytes of the emittedwake_arena_const_<memory>__blob.binsidecars. - Call
wake_model_init(&wake_ctx)andclf_model_init(&clf_ctx). Each one validates its own bind table, zeroes its persistent arena, hydrates any staged constant arena, then runs the per-operator init hooks. An unbound region fails here rather than at the first inference. - Run either context with
wake_model_runorclf_model_run. A run touches only the arenas bound for that prefix.
The persistent arenas have to stay separate, because they hold each model’s
state for the life of the program. The scratch arenas are the ones worth
sharing: one buffer, sized to the larger of wake_arena_<memory>_size and
clf_arena_<memory>_size and aligned to the larger of the two alignments, can
be bound into both.
Inspect the residency report to check the emitted plan.