Skip to content
heliaAOT
HELIA HUB

Generated files

The previous step wrote out/kws_ref/. Read this module before adding it to firmware: its headers describe the model I/O and memory contract, and its build files select the integration you requested.

The important parts of this walkthrough’s Zephyr output are:

selected Zephyr files, verified against module-layout.json
out/kws_ref/
├── README.md
├── LICENSE
├── aot_plan.json
├── aot_report.json
├── aot_residency.json
├── includes-api/
│ ├── aot_model.h
│ ├── aot_context.h
│ ├── aot_tensors.h
│ ├── aot_common.h
│ ├── aot_platform.h
│ └── aot_test_case.h
├── src/
│ ├── aot_model.c
│ ├── aot_context.c
│ ├── aot_tensors.c
│ ├── aot_constants.c
│ └── aot_test_case.c
└── zephyr/
├── module.yml
├── Kconfig
└── CMakeLists.txt

This is a selected file list: the full module also contains operator sources, headers and shared helpers. Per-node files identify the operator and its index. The generated README lists this model’s inputs, outputs, arenas and integration requirements. Its generic usage sketch is a starting point; the next page supplies the complete test application used in this walkthrough.

kws_ref is module.name, used for the directory and build target. aot is module.prefix, used for generated symbols and header names. Use distinct prefixes when linking multiple generated models.

The test files exist because test.enabled is true. The JSON report exists because memory.dump_residency_json is true. Other formats replace the Zephyr build files with their own entry points and adapt platform hooks; see the generated module reference.

After successful initialization, ctx.inputs[i] and ctx.outputs[i] describe model I/O in model order. Each descriptor provides data, size in bytes, zero_point and scale. Shape and element type come from the generated model summary and tensor declarations; an int8_t * storage pointer is not a promise that every model uses int8 elements.

For this KWS model, the input is 490 bytes and the output is 12 bytes. An application-facing wrapper can copy one already-prepared int8 input and return its int8 output:

#include <string.h>
#include "aot_model.h"
static aot_model_context_t kws_ctx = {0};
static int kws_ready = 0;
int32_t kws_start(void)
{
kws_ready = 0;
int32_t status = aot_model_init(&kws_ctx);
kws_ready = (status == 0);
return status;
}
int32_t kws_infer(const int8_t *input, size_t input_bytes,
int8_t *output, size_t output_bytes)
{
if (!kws_ready) {
return aot_status_not_initialized;
}
if (input == NULL || output == NULL ||
input_bytes != kws_ctx.inputs[0].size ||
output_bytes != kws_ctx.outputs[0].size) {
return -1; /* Application argument error. */
}
memcpy(kws_ctx.inputs[0].data, input, input_bytes);
int32_t status = aot_model_run(&kws_ctx);
if (status != 0) {
return status;
}
memcpy(output, kws_ctx.outputs[0].data, output_bytes);
return 0;
}

Call kws_start once and check for zero before calling kws_infer. The supplied golden test on the next page handles its own context and input copy, so you do not need this wrapper to run it.

aot_model_init initializes context, hydrates any staged constants and runs operator initialization. aot_model_run checks the lifecycle and executes the graph. Do not call aot_context_init separately after aot_model_init: it clears the lifecycle state. For internal recurrent state, initialize again when starting an independent sequence.

For a quantized value q, a descriptor’s scale and zero point give real_value = (q - zero_point) * scale. Application input preprocessing must produce the model’s expected layout and quantization; it is not supplied by aot_model_run.

The optional operator callback receives start/finish events with the table schedule. module.schedule: static omits those callbacks. Arena bindings and hydration state are module-global: serialize calls and do not assume two contexts provide independent concurrent instances.

aot_plan.json records, for every operator, the value each optimization knob resolved to, what was requested and by which rule, and the choices that are not knobs yet. aot_report.json explains it: the alternatives and whether they apply on this target, the measurements, and the facts behind the optimization hints in the results. Every conversion writes both; see Performance and accuracy options.

aot_residency.json reports scratch, persistent and constant arenas and tensor placements. Match its plan_hash to AOT_PLAN_HASH in aot_tensors.h before using sizes or region identifiers in an integration.

Role Meaning
Scratch Transient tensors whose storage can be reused according to their lifetimes
Persistent Writable state retained between runs and initialized on model init
Constant Weights and other fixed data; cold storage is read in place, staged data is copied to its runtime region

The walkthrough leaves memory.allocate_arenas enabled, so the module declares its buffers. Caller-owned arenas require aligned buffers, binding every region before init, and supplying any emitted constant sidecar data. Follow Memory before switching ownership.

A placement report describes the compiler’s plan. Check the firmware linker map to establish where the compiled buffers actually landed. The memory reports guide explains the report’s sizes and capacity fields.