Generated files
The previous step wrote out/kws_ref/. Read this module before adding it to
firmware: its headers describe the model I/O and memory contract, and its build
files select the integration you requested.
The tree
Section titled “The tree”The important parts of this walkthrough’s Zephyr output are:
out/kws_ref/├── README.md├── LICENSE├── aot_plan.json├── aot_report.json├── aot_residency.json├── includes-api/│ ├── aot_model.h│ ├── aot_context.h│ ├── aot_tensors.h│ ├── aot_common.h│ ├── aot_platform.h│ └── aot_test_case.h├── src/│ ├── aot_model.c│ ├── aot_context.c│ ├── aot_tensors.c│ ├── aot_constants.c│ └── aot_test_case.c└── zephyr/ ├── module.yml ├── Kconfig └── CMakeLists.txtThis is a selected file list: the full module also contains operator sources, headers and shared helpers. Per-node files identify the operator and its index. The generated README lists this model’s inputs, outputs, arenas and integration requirements. Its generic usage sketch is a starting point; the next page supplies the complete test application used in this walkthrough.
kws_ref is module.name, used for the directory and build target. aot is
module.prefix, used for generated symbols and header names. Use distinct
prefixes when linking multiple generated models.
The test files exist because test.enabled is true. The JSON report exists
because memory.dump_residency_json is true. Other formats replace the Zephyr
build files with their own entry points and adapt platform hooks; see the
generated module reference.
The public C API
Section titled “The public C API”After successful initialization, ctx.inputs[i] and ctx.outputs[i] describe
model I/O in model order. Each descriptor provides data, size in bytes,
zero_point and scale. Shape and element type come from the generated model
summary and tensor declarations; an int8_t * storage pointer is not a promise
that every model uses int8 elements.
For this KWS model, the input is 490 bytes and the output is 12 bytes. An application-facing wrapper can copy one already-prepared int8 input and return its int8 output:
#include <string.h>#include "aot_model.h"
static aot_model_context_t kws_ctx = {0};static int kws_ready = 0;
int32_t kws_start(void){ kws_ready = 0; int32_t status = aot_model_init(&kws_ctx); kws_ready = (status == 0); return status;}
int32_t kws_infer(const int8_t *input, size_t input_bytes, int8_t *output, size_t output_bytes){ if (!kws_ready) { return aot_status_not_initialized; } if (input == NULL || output == NULL || input_bytes != kws_ctx.inputs[0].size || output_bytes != kws_ctx.outputs[0].size) { return -1; /* Application argument error. */ } memcpy(kws_ctx.inputs[0].data, input, input_bytes); int32_t status = aot_model_run(&kws_ctx); if (status != 0) { return status; } memcpy(output, kws_ctx.outputs[0].data, output_bytes); return 0;}Call kws_start once and check for zero before calling kws_infer. The supplied
golden test on the next page handles its own context and input copy, so you do
not need this wrapper to run it.
aot_model_init initializes context, hydrates any staged constants and runs
operator initialization. aot_model_run checks the lifecycle and executes the
graph. Do not call aot_context_init separately after aot_model_init: it clears
the lifecycle state. For internal recurrent state, initialize again when
starting an independent sequence.
For a quantized value q, a descriptor’s scale and zero point give
real_value = (q - zero_point) * scale. Application input preprocessing must
produce the model’s expected layout and quantization; it is not supplied by
aot_model_run.
The optional operator callback receives start/finish events with the table
schedule. module.schedule: static omits those callbacks. Arena bindings and
hydration state are module-global: serialize calls and do not assume two
contexts provide independent concurrent instances.
The optimization plan and report
Section titled “The optimization plan and report”aot_plan.json records, for every operator, the value each optimization knob
resolved to, what was requested and by which rule, and the choices that are
not knobs yet. aot_report.json explains it: the alternatives and whether they
apply on this target, the measurements, and the facts behind the optimization
hints in the results. Every conversion writes both; see
Performance and accuracy options.
The residency report
Section titled “The residency report”aot_residency.json reports scratch, persistent and constant arenas and tensor
placements. Match its plan_hash to AOT_PLAN_HASH in aot_tensors.h before
using sizes or region identifiers in an integration.
| Role | Meaning |
|---|---|
| Scratch | Transient tensors whose storage can be reused according to their lifetimes |
| Persistent | Writable state retained between runs and initialized on model init |
| Constant | Weights and other fixed data; cold storage is read in place, staged data is copied to its runtime region |
The walkthrough leaves memory.allocate_arenas enabled, so the module declares
its buffers. Caller-owned arenas require aligned buffers, binding every region
before init, and supplying any emitted constant sidecar data. Follow
Memory before switching ownership.
A placement report describes the compiler’s plan. Check the firmware linker map to establish where the compiled buffers actually landed. The memory reports guide explains the report’s sizes and capacity fields.