Export and quantization
helia_edge.export converts a Keras model to LiteRT and returns the bytes with their identity. Conversion uses TensorFlow even when another part of your project uses the Torch backend. Run it in the appropriate environment and check the operators and input contract required by your deployment target.
Install the conversion capability
Section titled “Install the conversion capability”From the source checkout used by the installation guide:
uv sync --python 3.12 --extra litertUse a TensorFlow-backed Keras model. See Backend support for the supported Python environment.
Export and inspect
Section titled “Export and inspect”export_model takes an ExportSpec that states every setting that changes the exported bytes. With a trained model and a NumPy array representative_x matching its input shape:
from pathlib import Path
import numpy as npfrom helia_edge.export import ExportSpec, LiteRTRunner, export_model
spec = ExportSpec(precision="a8w8", io_dtype="int8", mode="concrete")result = export_model(model, spec, calibration=representative_x.astype(np.float32))print(result.sha256, result.inputs[0].scale, result.environment.helia_edge)predictions = LiteRTRunner(result.content).predict(representative_x)Path("model.tflite").write_bytes(result.content)| Precision | Graph | Input/output dtype |
|---|---|---|
fp32 |
float32 weights and compute | float32 |
fp16 |
native float16 graph | float16 |
a8w8 |
int8 activations and weights | int8 or float32 |
a16w8 |
int16 activations, int8 weights | int16 or float32 |
modeandio_dtypehave no defaults, and an invalid precision/dtype pair is rejected.a8w8anda16w8need calibration: a non-empty, finite float32 array shaped like the model input, used one sample at a time in stored order. Other precisions refuse calibration data.litertis the export format, and it needs the TensorFlow backend. A model with several inputs takes calibration as a mapping of input name to samples and thekerasorsaved_modelmode (see Streaming models). A model trained with the Torch backend is rebuilt from its params and weights in a process started withKERAS_BACKEND=tensorflow.- Choose representative inputs from the distribution the model will encounter. Compare predictions against the original model and evaluate task quality on held-out data before deployment.
ExportSpecvalues match helia-benchmark’s precision names.LiteRTRunnerruns model bytes with ai-edge-litert, rounding and saturating integer inputs. It runs native float16 graphs where the runtime has float16 kernels for every operator.
Export with a record
Section titled “Export with a record”export exports a model with its export record (helia-edge/export-record@1): one JSON file per artifact with the model spec, the weights digest, every export setting, the artifact’s sha256 and I/O, and the environment. It names the calibration and the weights by content hash only, never by path. write puts model.tflite, model.weights.h5 and record.json in a directory, and load_export_record rebuilds the Keras model from them, checking the weights against the digest:
from helia_edge.export import export, load_export_recordfrom helia_edge.models import ModelSpec, build, compact_tcn_params
spec = ModelSpec(params=compact_tcn_params(num_classes=2), input_shape=(256, 8))model = build(spec) # trained with a dynamic batchresult = export(model, precision="a8w8", io_dtype="int8", calibration=representative_x, spec=spec)record_path = result.write("exports/tcn-a8w8")model_again = load_export_record(record_path)- Static batch: the exported model has batch
batch_size(default 1). Withspec, the export isbuild(spec, batch_size=batch_size)with the model’s weights, whatever batch the model was built with, so a model trained with a dynamic batch exports with noSPACE_TO_BATCH_ND/BATCH_TO_SPACE_NDpairs or dynamic tensors, and the record re-exports the same bytes in another process. Withoutspec, the model is exported as it is and its batch must bebatch_size. - Weights digest: sha256 over each weight’s dtype, shape and value, in
model.weightsorder (not the weight names, which Keras numbers per process for unnamed layers). It is the same whether the weights come from.weights.h5,.kerasor an import (weights_importrecords the mapping and the source file’s sha256). - Calibrated precisions (
a8w8,a16w8) export withbatch_size1. Streamingresetsare increasing steps within the calibration. - Streaming models: pass the signal as
calibrationwith the stateresets;exportbuilds the state calibration withstream_calibration.with_golden(inputs, resets)adds ahelia-model-zoo/golden@2sequence run with LiteRT’s reference kernels. - Options:
ExportOptions(strict=True, state_tie_tolerance=0.01, mode="keras", dense_per_channel=True);mode="keras"keeps the model’s batch;concretetraces batch 1 and exportsbatch_size1 only. - Dense weights: for
a8w8anda16w8,dense_per_channel=FalsequantizesFULLY_CONNECTEDweights with one scale per tensor, while convolutions stay per channel. The CMSIS-NN int16FULLY_CONNECTEDkernel of LiteRT for Microcontrollers accepts per-tensor weights only, soa16w8exports for it useFalse; the reference and heliaRT kernels run either. The record names the setting, and float precisions refuseFalse.
Streaming models
Section titled “Streaming models”A streaming model processes one step per call and carries its recurrent state as explicit inputs and outputs. Name each state pair with state_input(k, ...) and state_output(k, ...): the runtime feeds state_out_k back as state_in_k and feeds zeros at the start of an independent sequence. StreamingLSTMCell is one LSTM step with the weight layout of keras.layers.LSTMCell.
import kerasimport numpy as npfrom helia_edge.export import ExportSpec, LiteRTStreamRunner, export_model, stream_calibrationfrom helia_edge.layers import StreamingLSTMCell, state_input, state_output
keras.utils.set_random_seed(0)x = keras.Input((16,), batch_size=1, name="features")h, c = state_input(0, (32,), batch_size=1), state_input(1, (32,), batch_size=1)h_next, c_next = StreamingLSTMCell(32)([x, h, c])prob = keras.layers.Dense(1, activation="sigmoid", name="prob")(h_next)model = keras.Model([x, h, c], [prob, state_output(0, h_next), state_output(1, c_next)])
features = np.random.default_rng(0).normal(size=(200, 16)).astype(np.float32)calibration = stream_calibration(model, {"features": features})result = export_model(model, ExportSpec(precision="a16w8", io_dtype="int16", mode="keras"), calibration)runner = LiteRTStreamRunner(result.content)fed, outputs = runner.run({"features": runner.encode("features", features[:50, None])})probabilities = runner.decode("prob", outputs["prob"])- Mode: use
keras.concreteconverts single-input models only, because its graph has no signature to name inputs and states.saved_modelexports a dynamic batch, so integer state pairs cannot be tied. - Records:
result.inputsandresult.outputslist tensors in the model’s subgraph order. State tensors have rolestateand the indexpairof their pair; the converter may reorder inputs, so address them by name. - Calibration:
stream_calibrationruns the Keras model over the signal inputs and returns every input, the state inputs filled with the states the model produces (zero at the start and at the steps inresets). - Tied state: with integer I/O (
a8w8withint8,a16w8withint16), the converter calibratesstate_in_kandstate_out_kseparately. Export then gives each pair one scale and zero point whose range covers both tensors’ ranges (forint16, the larger scale), so a rawstate_out_kkeeps its value as the nextstate_in_k; the bias of a fully connected or convolution layer that reads a state is requantized with it. It refuses when the tied scale differs from either original by more thanExportSpec.state_tie_toleranceof it (default 1%, at most 50%), or when an operator that needs the state tensor’s parameters unchanged (such as a concatenation, or a tanh writing the state) reads or writes it; a reshape passes the tied parameters on to its output. State inputs and outputs must pair up with equal shapes and types, and every output needs its own name (two outputs of one layer share it). The export record’siorecordsstate_scales_tied. - Running:
LiteRTStreamRunnerruns a stream on raw tensors, carrying the state unchanged; it returns every input as fed and every output, stacked by step.LiteRTRunnerstays single-input. exporttakes streaming models with one signal input; see Export with a record.
The command line
Section titled “The command line”helia-edge export create builds a model from a ModelSpec file (YAML or JSON), loads its weights, exports it with export and writes the record directory. helia-edge export reproduce exports a record’s model again and compares:
KERAS_BACKEND=tensorflow helia-edge export create spec.yaml --weights trained.weights.h5 \ --precision a8w8 --calibration calibration.npy --out exports/tcn-a8w8KERAS_BACKEND=tensorflow helia-edge export reproduce exports/tcn-a8w8/record.json \ --weights exports/tcn-a8w8/model.weights.h5 --calibration calibration.npy-
create:--weightsis a Keras weights file (.weights.h5or.keras), or with--mapping NAMEthe source file imported through that mapping of the spec’s family (see Import weights).--io-dtypedefaults tofloat32,float16,int8orint16forfp32,fp16,a8w8anda16w8.--calibrationand--golden-inputsare.npyfiles, with--resetsand--golden-resetsrepeated for each reset step.--batch-size(default 1),--mode keras|concrete,--no-strict,--state-tie-toleranceand--dense-per-tensorare passed toexport. A missing argument file or a--batch-sizeoutside 1 to 2³¹−1 is a usage error (exit 2). A refused export (an invalid spec, setting or input file, a backend other than TensorFlow) exits 1 with an error and writes no record. -
reproduce:--weightsis the record’s weights (any file with the recorded digest), or for imported weights the source file with the recorded sha256.--calibrationis required when the record has calibration and refused when it has none;--golden-inputsregenerates and compares the golden (without it, the output says the golden was not compared). Amodel.tfliteorgolden.npzbeside the record is compared with the record’s size and sha256. Exit codes:- 0: the weights digest, the export settings, the artifact’s sha256 and size, the I/O, the golden (with
--golden-inputs) and the files beside the record match; - 1: something differs, and each difference is listed;
- 2: the environment differs from the recorded one, and nothing is exported.
--allow-env-mismatchcompares anyway and lists the differences. Also when this process cannot export: Keras is missing or its backend is not TensorFlow; - 3: the record is missing, invalid or has no model spec, or an input file is missing, unreadable, does not match its recorded hash or is not used by the record;
- 4: the export itself failed with an error, printed with its traceback; nothing was compared.
- 0: the weights digest, the export settings, the artifact’s sha256 and size, the I/O, the golden (with
-
schema: prints the record’s JSON Schema. -
Environment: the record names the helia-edge version, install source and commit, Python, the platform and the NumPy, Keras, TensorFlow and ai-edge-litert versions. Other versions may export other bytes.
-
Provenance:
environment.helia_edge.sourcesays whether the version identifies the code that exported:release: installed from a package index (a wheel built from a modified tree and served from a private index looks the same, so prefervcsfor published exports);vcs: installed from a git URL, at the recorded commit (from a localgit+file://repository, that commit may never have been pushed);local: installed from a directory or an archive (local or by URL), or with only egg-info metadata (a source tree, some system packages), none of which the version identifies;unknown.
createwarns unless it isreleaseorvcs; with--require-provenanceit refuses. For exports you publish, install from git at a commit:uv pip install 'helia-edge @ git+https://github.com/AmbiqAI/helia-edge@<commit>'. -
Layer names: unnamed layers’ names become tensor names in the exported bytes, and Keras numbers them per process.
exportwithspec,load_export_recordand both commands build with the numbering started from zero, through a private Keras name table. If a Keras release no longer numbers names through it, they warn (RuntimeWarning) that the bytes may depend on layers built earlier in the process.
Hand off to an Ambiq deployment
Section titled “Hand off to an Ambiq deployment”The exported model is one part of the deployment. Select the inference runtime and check its operator and precision support before integrating it into firmware. Training backend support does not imply that every model can run on every target.
| Keep with the model | Why the deployment needs it |
|---|---|
| Input shape, layout and dtype | Defines how firmware supplies each inference window. |
| Preprocessing parameters | Keeps scaling, filtering and feature extraction consistent with training. |
| Output names and interpretation | Defines class order, thresholds or reconstruction units. |
| Calibration and quantization settings | Explains how the converted artifact was produced. |
| Representative inputs and expected outputs | Enables comparison between host inference and the deployed runtime. |
| Model and source revisions | Identifies the exact artifact being evaluated. |
Validate task quality after conversion, then measure latency and memory with the model on your intended Ambiq board. Host conversion and operation counts do not establish device performance. For streaming models, include the state shapes, initialization and update order described in the model’s guide.