Skip to content
heliaEDGE
User guide
HELIA

Export and quantization

helia_edge.export converts a Keras model to LiteRT and returns the bytes with their identity. Conversion uses TensorFlow even when another part of your project uses the Torch backend. Run it in the appropriate environment and check the operators and input contract required by your deployment target.

From the source checkout used by the installation guide:

Terminal window
uv sync --python 3.12 --extra litert

Use a TensorFlow-backed Keras model. See Backend support for the supported Python environment.

export_model takes an ExportSpec that states every setting that changes the exported bytes. With a trained model and a NumPy array representative_x matching its input shape:

from pathlib import Path
import numpy as np
from helia_edge.export import ExportSpec, LiteRTRunner, export_model
spec = ExportSpec(precision="a8w8", io_dtype="int8", mode="concrete")
result = export_model(model, spec, calibration=representative_x.astype(np.float32))
print(result.sha256, result.inputs[0].scale, result.environment.helia_edge)
predictions = LiteRTRunner(result.content).predict(representative_x)
Path("model.tflite").write_bytes(result.content)
Precision Graph Input/output dtype
fp32 float32 weights and compute float32
fp16 native float16 graph float16
a8w8 int8 activations and weights int8 or float32
a16w8 int16 activations, int8 weights int16 or float32
  • mode and io_dtype have no defaults, and an invalid precision/dtype pair is rejected.
  • a8w8 and a16w8 need calibration: a non-empty, finite float32 array shaped like the model input, used one sample at a time in stored order. Other precisions refuse calibration data.
  • litert is the export format, and it needs the TensorFlow backend. A model with several inputs takes calibration as a mapping of input name to samples and the keras or saved_model mode (see Streaming models). A model trained with the Torch backend is rebuilt from its params and weights in a process started with KERAS_BACKEND=tensorflow.
  • Choose representative inputs from the distribution the model will encounter. Compare predictions against the original model and evaluate task quality on held-out data before deployment.
  • ExportSpec values match helia-benchmark’s precision names.
  • LiteRTRunner runs model bytes with ai-edge-litert, rounding and saturating integer inputs. It runs native float16 graphs where the runtime has float16 kernels for every operator.

export exports a model with its export record (helia-edge/export-record@1): one JSON file per artifact with the model spec, the weights digest, every export setting, the artifact’s sha256 and I/O, and the environment. It names the calibration and the weights by content hash only, never by path. write puts model.tflite, model.weights.h5 and record.json in a directory, and load_export_record rebuilds the Keras model from them, checking the weights against the digest:

from helia_edge.export import export, load_export_record
from helia_edge.models import ModelSpec, build, compact_tcn_params
spec = ModelSpec(params=compact_tcn_params(num_classes=2), input_shape=(256, 8))
model = build(spec) # trained with a dynamic batch
result = export(model, precision="a8w8", io_dtype="int8", calibration=representative_x, spec=spec)
record_path = result.write("exports/tcn-a8w8")
model_again = load_export_record(record_path)
  • Static batch: the exported model has batch batch_size (default 1). With spec, the export is build(spec, batch_size=batch_size) with the model’s weights, whatever batch the model was built with, so a model trained with a dynamic batch exports with no SPACE_TO_BATCH_ND/BATCH_TO_SPACE_ND pairs or dynamic tensors, and the record re-exports the same bytes in another process. Without spec, the model is exported as it is and its batch must be batch_size.
  • Weights digest: sha256 over each weight’s dtype, shape and value, in model.weights order (not the weight names, which Keras numbers per process for unnamed layers). It is the same whether the weights come from .weights.h5, .keras or an import (weights_import records the mapping and the source file’s sha256).
  • Calibrated precisions (a8w8, a16w8) export with batch_size 1. Streaming resets are increasing steps within the calibration.
  • Streaming models: pass the signal as calibration with the state resets; export builds the state calibration with stream_calibration. with_golden(inputs, resets) adds a helia-model-zoo/golden@2 sequence run with LiteRT’s reference kernels.
  • Options: ExportOptions(strict=True, state_tie_tolerance=0.01, mode="keras", dense_per_channel=True); mode="keras" keeps the model’s batch; concrete traces batch 1 and exports batch_size 1 only.
  • Dense weights: for a8w8 and a16w8, dense_per_channel=False quantizes FULLY_CONNECTED weights with one scale per tensor, while convolutions stay per channel. The CMSIS-NN int16 FULLY_CONNECTED kernel of LiteRT for Microcontrollers accepts per-tensor weights only, so a16w8 exports for it use False; the reference and heliaRT kernels run either. The record names the setting, and float precisions refuse False.

A streaming model processes one step per call and carries its recurrent state as explicit inputs and outputs. Name each state pair with state_input(k, ...) and state_output(k, ...): the runtime feeds state_out_k back as state_in_k and feeds zeros at the start of an independent sequence. StreamingLSTMCell is one LSTM step with the weight layout of keras.layers.LSTMCell.

import keras
import numpy as np
from helia_edge.export import ExportSpec, LiteRTStreamRunner, export_model, stream_calibration
from helia_edge.layers import StreamingLSTMCell, state_input, state_output
keras.utils.set_random_seed(0)
x = keras.Input((16,), batch_size=1, name="features")
h, c = state_input(0, (32,), batch_size=1), state_input(1, (32,), batch_size=1)
h_next, c_next = StreamingLSTMCell(32)([x, h, c])
prob = keras.layers.Dense(1, activation="sigmoid", name="prob")(h_next)
model = keras.Model([x, h, c], [prob, state_output(0, h_next), state_output(1, c_next)])
features = np.random.default_rng(0).normal(size=(200, 16)).astype(np.float32)
calibration = stream_calibration(model, {"features": features})
result = export_model(model, ExportSpec(precision="a16w8", io_dtype="int16", mode="keras"), calibration)
runner = LiteRTStreamRunner(result.content)
fed, outputs = runner.run({"features": runner.encode("features", features[:50, None])})
probabilities = runner.decode("prob", outputs["prob"])
  • Mode: use keras. concrete converts single-input models only, because its graph has no signature to name inputs and states. saved_model exports a dynamic batch, so integer state pairs cannot be tied.
  • Records: result.inputs and result.outputs list tensors in the model’s subgraph order. State tensors have role state and the index pair of their pair; the converter may reorder inputs, so address them by name.
  • Calibration: stream_calibration runs the Keras model over the signal inputs and returns every input, the state inputs filled with the states the model produces (zero at the start and at the steps in resets).
  • Tied state: with integer I/O (a8w8 with int8, a16w8 with int16), the converter calibrates state_in_k and state_out_k separately. Export then gives each pair one scale and zero point whose range covers both tensors’ ranges (for int16, the larger scale), so a raw state_out_k keeps its value as the next state_in_k; the bias of a fully connected or convolution layer that reads a state is requantized with it. It refuses when the tied scale differs from either original by more than ExportSpec.state_tie_tolerance of it (default 1%, at most 50%), or when an operator that needs the state tensor’s parameters unchanged (such as a concatenation, or a tanh writing the state) reads or writes it; a reshape passes the tied parameters on to its output. State inputs and outputs must pair up with equal shapes and types, and every output needs its own name (two outputs of one layer share it). The export record’s io records state_scales_tied.
  • Running: LiteRTStreamRunner runs a stream on raw tensors, carrying the state unchanged; it returns every input as fed and every output, stacked by step. LiteRTRunner stays single-input.
  • export takes streaming models with one signal input; see Export with a record.

helia-edge export create builds a model from a ModelSpec file (YAML or JSON), loads its weights, exports it with export and writes the record directory. helia-edge export reproduce exports a record’s model again and compares:

Terminal window
KERAS_BACKEND=tensorflow helia-edge export create spec.yaml --weights trained.weights.h5 \
--precision a8w8 --calibration calibration.npy --out exports/tcn-a8w8
KERAS_BACKEND=tensorflow helia-edge export reproduce exports/tcn-a8w8/record.json \
--weights exports/tcn-a8w8/model.weights.h5 --calibration calibration.npy
  • create: --weights is a Keras weights file (.weights.h5 or .keras), or with --mapping NAME the source file imported through that mapping of the spec’s family (see Import weights). --io-dtype defaults to float32, float16, int8 or int16 for fp32, fp16, a8w8 and a16w8. --calibration and --golden-inputs are .npy files, with --resets and --golden-resets repeated for each reset step. --batch-size (default 1), --mode keras|concrete, --no-strict, --state-tie-tolerance and --dense-per-tensor are passed to export. A missing argument file or a --batch-size outside 1 to 2³¹−1 is a usage error (exit 2). A refused export (an invalid spec, setting or input file, a backend other than TensorFlow) exits 1 with an error and writes no record.

  • reproduce: --weights is the record’s weights (any file with the recorded digest), or for imported weights the source file with the recorded sha256. --calibration is required when the record has calibration and refused when it has none; --golden-inputs regenerates and compares the golden (without it, the output says the golden was not compared). A model.tflite or golden.npz beside the record is compared with the record’s size and sha256. Exit codes:

    • 0: the weights digest, the export settings, the artifact’s sha256 and size, the I/O, the golden (with --golden-inputs) and the files beside the record match;
    • 1: something differs, and each difference is listed;
    • 2: the environment differs from the recorded one, and nothing is exported. --allow-env-mismatch compares anyway and lists the differences. Also when this process cannot export: Keras is missing or its backend is not TensorFlow;
    • 3: the record is missing, invalid or has no model spec, or an input file is missing, unreadable, does not match its recorded hash or is not used by the record;
    • 4: the export itself failed with an error, printed with its traceback; nothing was compared.
  • schema: prints the record’s JSON Schema.

  • Environment: the record names the helia-edge version, install source and commit, Python, the platform and the NumPy, Keras, TensorFlow and ai-edge-litert versions. Other versions may export other bytes.

  • Provenance: environment.helia_edge.source says whether the version identifies the code that exported:

    • release: installed from a package index (a wheel built from a modified tree and served from a private index looks the same, so prefer vcs for published exports);
    • vcs: installed from a git URL, at the recorded commit (from a local git+file:// repository, that commit may never have been pushed);
    • local: installed from a directory or an archive (local or by URL), or with only egg-info metadata (a source tree, some system packages), none of which the version identifies;
    • unknown.

    create warns unless it is release or vcs; with --require-provenance it refuses. For exports you publish, install from git at a commit: uv pip install 'helia-edge @ git+https://github.com/AmbiqAI/helia-edge@<commit>'.

  • Layer names: unnamed layers’ names become tensor names in the exported bytes, and Keras numbers them per process. export with spec, load_export_record and both commands build with the numbering started from zero, through a private Keras name table. If a Keras release no longer numbers names through it, they warn (RuntimeWarning) that the bytes may depend on layers built earlier in the process.

The exported model is one part of the deployment. Select the inference runtime and check its operator and precision support before integrating it into firmware. Training backend support does not imply that every model can run on every target.

Keep with the model Why the deployment needs it
Input shape, layout and dtype Defines how firmware supplies each inference window.
Preprocessing parameters Keeps scaling, filtering and feature extraction consistent with training.
Output names and interpretation Defines class order, thresholds or reconstruction units.
Calibration and quantization settings Explains how the converted artifact was produced.
Representative inputs and expected outputs Enables comparison between host inference and the deployed runtime.
Model and source revisions Identifies the exact artifact being evaluated.

Validate task quality after conversion, then measure latency and memory with the model on your intended Ambiq board. Host conversion and operation counts do not establish device performance. For streaming models, include the state shapes, initialization and update order described in the model’s guide.