# Generated files

The previous step wrote `out/kws_ref/`. Read this module before adding it to
firmware: its headers describe the model I/O and memory contract, and its build
files select the integration you requested.

## The tree

The important parts of this walkthrough's Zephyr output are:

```text title="selected Zephyr files, verified against module-layout.json"
out/kws_ref/
├── README.md
├── LICENSE
├── aot_plan.json
├── aot_report.json
├── aot_residency.json
├── includes-api/
│   ├── aot_model.h
│   ├── aot_context.h
│   ├── aot_tensors.h
│   ├── aot_common.h
│   ├── aot_platform.h
│   └── aot_test_case.h
├── src/
│   ├── aot_model.c
│   ├── aot_context.c
│   ├── aot_tensors.c
│   ├── aot_constants.c
│   └── aot_test_case.c
└── zephyr/
    ├── module.yml
    ├── Kconfig
    └── CMakeLists.txt
```

This is a selected file list: the full module also contains operator sources,
headers and shared helpers. Per-node files identify the operator and its index.
The generated README lists this model's inputs, outputs, arenas and integration
requirements. Its generic usage sketch is a starting point; the next page
supplies the complete test application used in this walkthrough.

`kws_ref` is `module.name`, used for the directory and build target. `aot` is
`module.prefix`, used for generated symbols and header names. Use distinct
prefixes when linking multiple generated models.

The test files exist because `test.enabled` is true. The JSON report exists
because `memory.dump_residency_json` is true. Other formats replace the Zephyr
build files with their own entry points and adapt platform hooks; see the
[generated module reference](https://ambiqai.github.io/helia-aot/reference/module/).

## The public C API

After successful initialization, `ctx.inputs[i]` and `ctx.outputs[i]` describe
model I/O in model order. Each descriptor provides `data`, `size` **in bytes**,
`zero_point` and `scale`. Shape and element type come from the generated model
summary and tensor declarations; an `int8_t *` storage pointer is not a promise
that every model uses int8 elements.

For this KWS model, the input is 490 bytes and the output is 12 bytes. An
application-facing wrapper can copy one already-prepared int8 input and return
its int8 output:

```c
#include <string.h>
#include "aot_model.h"

static aot_model_context_t kws_ctx = {0};
static int kws_ready = 0;

int32_t kws_start(void)
{
    kws_ready = 0;
    int32_t status = aot_model_init(&kws_ctx);
    kws_ready = (status == 0);
    return status;
}

int32_t kws_infer(const int8_t *input, size_t input_bytes,
                  int8_t *output, size_t output_bytes)
{
    if (!kws_ready) {
        return aot_status_not_initialized;
    }
    if (input == NULL || output == NULL ||
        input_bytes != kws_ctx.inputs[0].size ||
        output_bytes != kws_ctx.outputs[0].size) {
        return -1; /* Application argument error. */
    }
    memcpy(kws_ctx.inputs[0].data, input, input_bytes);
    int32_t status = aot_model_run(&kws_ctx);
    if (status != 0) {
        return status;
    }
    memcpy(output, kws_ctx.outputs[0].data, output_bytes);
    return 0;
}
```

Call `kws_start` once and check for zero before calling `kws_infer`. The supplied
golden test on the next page handles its own context and input copy, so you do
not need this wrapper to run it.

`aot_model_init` initializes context, hydrates any staged constants and runs
operator initialization. `aot_model_run` checks the lifecycle and executes the
graph. Do not call `aot_context_init` separately after `aot_model_init`: it clears
the lifecycle state. For internal recurrent state, initialize again when
starting an independent sequence.

For a quantized value `q`, a descriptor's scale and zero point give
`real_value = (q - zero_point) * scale`. Application input preprocessing must
produce the model's expected layout and quantization; it is not supplied by
`aot_model_run`.

The optional operator callback receives start/finish events with the table
schedule. `module.schedule: static` omits those callbacks. Arena bindings and
hydration state are module-global: serialize calls and do not assume two
contexts provide independent concurrent instances.

## The optimization plan and report

`aot_plan.json` records, for every operator, the value each optimization knob
resolved to, what was requested and by which rule, and the choices that are
not knobs yet. `aot_report.json` explains it: the alternatives and whether they
apply on this target, the measurements, and the facts behind the optimization
hints in the results. Every conversion writes both; see
[Performance and accuracy options](https://ambiqai.github.io/helia-aot/guide/options/#the-plan-and-the-report).

## The residency report

`aot_residency.json` reports scratch, persistent and constant arenas and tensor
placements. Match its `plan_hash` to `AOT_PLAN_HASH` in `aot_tensors.h` before
using sizes or region identifiers in an integration.

| Role | Meaning |
| --- | --- |
| Scratch | Transient tensors whose storage can be reused according to their lifetimes |
| Persistent | Writable state retained between runs and initialized on model init |
| Constant | Weights and other fixed data; cold storage is read in place, staged data is copied to its runtime region |

The walkthrough leaves `memory.allocate_arenas` enabled, so the module declares
its buffers. Caller-owned arenas require aligned buffers, binding every region
before init, and supplying any emitted constant sidecar data. Follow
[Memory](https://ambiqai.github.io/helia-aot/guide/memory/) before switching ownership.

A placement report describes the compiler's plan. Check the firmware linker map
to establish where the compiled buffers actually landed. The
[memory reports guide](https://ambiqai.github.io/helia-aot/guide/memory-reports/) explains the report's
sizes and capacity fields.
