Skip to content
heliaAOT
HELIA HUB

Static schedule

Convert one model under both dispatch modes and diff the generated aot_model.c. The table mode walks a function-pointer table; the static mode emits the operator sequence as direct calls.

Field Value
Model kws_ref, MLPerf Tiny keyword spotting, int8
Target apollo510_evb
Shows module.schedule, generated dispatch code and callback availability
CI Conversion and host compilation are declared in the examples CI job

The model is audio/mlperf-tiny/kws_ref/model.tflite from helia-model-zoo: int8, 53,936 bytes, SHA-256 aeea4368…bd0ae, fetched by run.sh and checked against the fixed SHA-256 recorded in the checked-in run script. Model licensing is not established by these hashes; see Model provenance.

Use a repository checkout with Bash and the environment from Running examples: uv sync --frozen --group ci, then activate .venv. Fetched models require network access, curl and sha256sum or shasum. Run the commands below from this example directory. Conversion runs on the host; it does not execute the emitted firmware.

Terminal window
./run.sh
diff out/kws_ref_table/src/aot_model.c out/kws_ref_static/src/aot_model.c

config-table.yaml and config-static.yaml change the execution schedule and use different module names to retain both outputs. The behavior-affecting choice is:

module:
schedule: static # or table, the default

The table build declares one rodata entry per operator and a loop over it:

typedef int32_t (*aot_op_fn)(aot_model_context_t *ctx);
typedef struct {
aot_op_fn init; // Per-operator init hook
aot_op_fn run; // Per-operator run hook (resolves its own I/O)
int32_t id; // Stable AIR operator id, forwarded to callbacks
} aot_op_entry_t;
static const aot_op_entry_t aot_op_table[13] = {
{ aot_conv_2d_0_init, aot_conv_2d_0_run, 0 },
{ aot_depthwise_conv_2d_1_init, aot_depthwise_conv_2d_1_run, 1 },
...
};

The static build has no table and no loop. Every call target is visible to the linker and the optimizer, which is what makes inlining and tail calls available across the sequence.

The trade is the per-node callback seam. The table loop is where an RTOS can yield between operators and where per-layer counters are sampled, through ctx->callback. The static build omits it. Both modes emit the same public API, the same status codes, the same lifecycle latch and the same constant hydration, so the choice is dispatch only and can be made per build.

Source line counts do not measure firmware size or inference time. The size and cycle effect belongs to a real build and a real board: measure it with Profile with hpx rather than inferring it from the source.