Skip to content
heliaAOT
HELIA HUB

defines

Memory planning data model.

Describes what a planner consumes and produces: arena roles and usage, tensor bindings, lifetimes and allocations, the placement constraints read from configuration, and the resulting :class:MemoryPlan.

Machine-readable model

class

ArenaRole

Python

Logical role of a memory arena.

helia_aot/memory/defines.py:18

ArenaRole()

Logical role of a memory arena.

Every tensor in a :class:MemoryPlan lives in an arena tagged with one of these roles. Roles partition the planner output so scratch reuse, persistent zero-init, and constant residency can be reasoned about independently.

attribute

constant

Python

helia_aot/memory/defines.py:40

constant = auto()

Read-only weight storage. Can be cold (the arena buffer is the cold-storage blob and kernels read it in place) or staged (the arena buffer is a writable runtime copy hydrated from a separate source blob).

class

Planner's emit intent for a single tensor's storage.

helia_aot/memory/defines.py:43

TensorBinding()

Planner’s emit intent for a single tensor’s storage.

Every tensor in a :class:MemoryPlan is bound to an arena slot at (role, memory, offset). There are no per-tensor C symbols anymore — scratch, persistent, and constant tensors are all descriptors against arena buffers exposed via arena_buffers[region].

Constant arenas come in two shapes, distinguished purely by whether source_memory matches memory:

  • Cold (source_memory == memory): the arena buffer itself is a static const blob in cold storage. Kernels read it directly; no hydration is required.
  • Staged (source_memory != memory): the arena buffer is a writable runtime copy in memory; a separate read-only source blob lives in source_memory and the caller (or the weak <prefix>_hydrate_constants helper) copies it in before the first model_run.

Scratch and persistent bindings always have source_memory == memory (they are writable arenas with no cold-storage source).

attribute

memory

Python

helia_aot/memory/defines.py:81

memory: MemoryType = Field(..., description='Runtime memory of the arena slot.')

Runtime (kernel-visible) memory of the arena slot.

attribute

helia_aot/memory/defines.py:82

source_memory: MemoryType = Field(..., description='Cold-storage memory the slot is sourced from.')

Cold-storage memory the slot is sourced from. Equals memory for scratch, persistent, and cold constants; differs from memory for staged constants.

class

Tensor operator lifetime.

helia_aot/memory/defines.py:86

TensorLifetime()

Tensor operator lifetime.

attribute

start_op

Python

helia_aot/memory/defines.py:97

start_op: int = Field(..., description='Index of the first operator that uses/defines it')

Index of the first operator that uses/defines it

method

add_op

Python

Add an operator index to the lifetime.

helia_aot/memory/defines.py:100

add_op(op_idx: int)

Add an operator index to the lifetime.

Parameters of add_op
NameTypeDefaultDescription
op_idxintRequiredOperator index to add
method

merge

Python

Merge another lifetime into this one.

helia_aot/memory/defines.py:109

merge(other: TensorLifetime)

Merge another lifetime into this one.

This updates the start and end operators to encompass both lifetimes. Useful when tensors have aliases.

Parameters of merge
NameTypeDefaultDescription
otherTensorLifetimeRequiredAnother tensor lifetime to merge
class

Metadata for a tensor allocation.

helia_aot/memory/defines.py:126

TensorAllocation()

Metadata for a tensor allocation.

attribute

memory

Python

helia_aot/memory/defines.py:147

memory: MemoryType = Field(..., description='Memory type for the allocation')

Memory type for the allocation (runtime / kernel-visible memory of the arena slot)

attribute

binding

Python

helia_aot/memory/defines.py:150

binding: TensorBinding | None = Field(default=None, description="Planner-assigned binding describing how this tensor's storage is materialized.")

Planner’s binding describing (role, memory, offset, source_memory) for the slot. Defaults to None for callers that construct :class:TensorAllocation directly; the bundled :class:GreedyMemoryPlanner always populates it. Emit templates dispatch on binding.role and on the binding.source_memory == binding.memory predicate (cold vs staged constant).

class

Metadata for a memory arena.

helia_aot/memory/defines.py:156

ArenaUsage()

Metadata for a memory arena.

attribute

used

Python

helia_aot/memory/defines.py:179

used: int = Field(default=0, description='Bytes actually used')

Bytes actually used. For bump-allocated arenas (constant, persistent) this equals total_size.

attribute

role

Python

helia_aot/memory/defines.py:180

role: ArenaRole = Field(default=ArenaRole.scratch, description='Logical role of the arena (scratch | persistent | constant).')

Logical role of the arena.

attribute

helia_aot/memory/defines.py:184

source_memory: MemoryType = Field(default=None, description="Cold-storage source memory for the arena's contents. Equals ``memory`` for scratch, persistent, and cold constant arenas; differs from ``memory`` for staged constant arenas where bytes are copied at boot from this source memory into the writable runtime arena. Defaults to None for backward compatibility with callers that construct ArenaUsage directly; the bundled GreedyPlanner always populates it.")

The cold-storage memory the arena’s contents are sourced from. For scratch and persistent arenas this equals memory (purely runtime, no source blob). For constant arenas it equals memory in the cold case (arena buffer is the cold blob in place) and differs from memory in the staged case (arena buffer is a writable runtime copy of a source blob in source_memory).

attribute

alignment

Python

helia_aot/memory/defines.py:197

alignment: int = Field(default=16, description='Resolved alignment of the arena base symbol (and floor for slot offsets within it). Stamped by the planner from max(implementation floor, platform.min_alignment, MemoryConstraint.arena_alignment).')
attribute

is_staged

Python

True iff this arena has a distinct cold source memory.

helia_aot/memory/defines.py:216

is_staged: bool

True iff this arena has a distinct cold source memory.

Equivalent to source_memory is not None and source_memory != memory. Centralizes the cold-vs-staged predicate so templates and handlers do not duplicate it.

For scratch and persistent arenas this is always False (they have no cold source). For constant arenas it distinguishes the two emit shapes:

  • False (cold): the arena buffer itself is a static const blob in cold storage.
  • True (staged): the arena buffer is a writable runtime copy hydrated from a separate source blob in source_memory.
class

User-defined memory constraint

helia_aot/memory/defines.py:236

MemoryConstraint()

User-defined memory constraint

attribute

max_size

Python

helia_aot/memory/defines.py:260

max_size: int | None = Field(default=None, description='Maximum size in bytes, or None for no limit')

Maximum size in bytes, or None for no limit

attribute

helia_aot/memory/defines.py:261

arena_alignment: int | None = Field(default=None, description='Optional per-arena alignment floor in bytes. When set, applied to both the arena base symbol and every slot. When None, only the 16-byte implementation floor applies to the base; per-slot alignment is driven by platform / dtype / tensor hints.')

Optional per-arena alignment floor in bytes. When set, the planner uses max(arena_alignment, platform.min_alignment) for both the arena base symbol and the per-slot offset of every tensor packed into it. When None (the default), only the implementation floor (16 bytes) is applied to the arena base symbol; per-slot alignment continues to use max(platform.min_alignment, dtype_alignment_floor, tensor.alignment_hint) so dtype-natural and kernel-driven alignment are still honored. Bump arena_alignment above the default for memories that back DMA-driven hydration paths needing stronger alignment than the platform’s MVE/Helium floor (e.g. cacheline-sized PSRAM transfers).

class

Top-level memory plan for tensors.

helia_aot/memory/defines.py:293

MemoryPlan()

Top-level memory plan for tensors.

Every tensor allocation is a slot in some arena. Three arena maps partition storage by role:

  • :attr:arena_usages — writable scratch arenas (one per writable memory bank).

  • :attr:persistent_arenas — writable persistent arenas (one per writable memory bank that received a persistent tensor). Separate from scratch so the scratch arena stays purely transient and may be aliased across models.

  • :attr:constant_arenas — constant arenas. Two shapes share the same map:

    • Cold: arena.source_memory == arena.memory. The arena buffer itself is the read-only blob in cold storage; kernels read it directly. No hydration required.
    • Staged: arena.source_memory != arena.memory. The arena buffer is a writable runtime copy; a separate source blob in arena.source_memory is hydrated into it before the first model_run.
attribute

helia_aot/memory/defines.py:335

tensor_allocs: dict[str, TensorAllocation] = Field(..., default_factory=dict, description='Tensor allocations')

Tensor allocations. Each carries a :class:TensorBinding recording (role, memory, offset, source_memory).

attribute

helia_aot/memory/defines.py:337

constant_arenas: dict[MemoryType, ArenaUsage] = Field(..., default_factory=dict, description='Per-bank constant arenas (one per runtime memory).')

Constant arenas, keyed by runtime (destination) MemoryType. Per-arena single-source-memory invariant holds (every constant in the same arena has the same source_memory) so cold arenas and staged arenas are both guaranteed to be a single contiguous source blob.

attribute

helia_aot/memory/defines.py:342

persistent_arenas: dict[MemoryType, ArenaUsage] = Field(..., default_factory=dict, description='Per-bank persistent (resource-variable) arenas.')

Persistent arenas, keyed by writable MemoryType.

method

Get the allocation metadata for a given tensor ID.

helia_aot/memory/defines.py:349

get_allocation(tensor_id: str) -> TensorAllocation

Get the allocation metadata for a given tensor ID.

Parameters of get_allocation
NameTypeDefaultDescription
tensor_idstrRequiredThe tensor ID to look up.
Returns of get_allocation
ValueTypeDescription
TensorAllocationTensorAllocationThe allocation metadata for the tensor.
Errors raised by get_allocation
TypeDescription
KeyErrorIf the tensor ID is not found in the allocations.
class

Memory planner type

helia_aot/memory/defines.py:364

MemoryPlannerType()

Memory planner type

Selecting an experimental planner is an explicit act: the default layout of an existing conversion is unchanged until one is opted into.

All planners honor user-specified memory-region placement: per-tensor memory pins and constant_destination_memory routing from memory.tensors rules, and per-memory max_size budgets from memory.constraints. The rules are identical; the outcomes are not necessarily. Unpinned scratch spills to the next writable bank when the current one fills, so a different visit order can move a tensor between banks — an access-latency change, not only a footprint one.

attribute

helia_aot/memory/defines.py:395

hill_climb: str = auto()

Experimental, opt-in local search over greedy-by-size placement orders; its scratch high-water mark is never worse than greedy_by_size. Tunable via memory.planner_options (iterations, seed, max_stall_iterations).

class

This class provides the baseline set of tensor attribute rules.

helia_aot/memory/defines.py:398

TensorAttributes()

This class provides the baseline set of tensor attribute rules.

attribute

memory

Python

helia_aot/memory/defines.py:421

memory: MemoryType = Field(default=MemoryType.DTCM, description='Memory placement for tensors')

Memory placement for tensors. For constants this is the source (cold-storage) memory where the tensor’s bytes live in the image. When constant_destination_memory is set and differs from memory, the runtime arena lives in that destination memory and a hydration copy is required; otherwise the arena lives in memory itself and is read-only (cold).

attribute

helia_aot/memory/defines.py:422

constant_destination_memory: MemoryType | None = Field(default=None, description='Per-tensor runtime destination memory for constants. When None (default) the constant is read in place from its source memory; when set the constant is hydrated into a writable arena slot in the destination memory.')

Override the runtime (kernel-visible) memory for a constant. Required only when the runtime memory must differ from the source memory (e.g. weights stored in MRAM but read from DTCM/SRAM at runtime). When None (default), the runtime memory equals memory and no hydration step is needed. Must reference a writable memory of the target platform when set.