# Export and quantization

`helia_edge.export` converts a Keras model to LiteRT and returns the bytes with their identity. Conversion uses TensorFlow even when another part of your project uses the Torch backend. Run it in the appropriate environment and check the operators and input contract required by your deployment target.

## Install the conversion capability

From the source checkout used by the installation guide:

```sh
uv sync --python 3.12 --extra litert
```

Use a TensorFlow-backed Keras model. See [Backend support](https://ambiqai.github.io/helia-edge/getting-started/backends/) for the supported Python environment.

## Export and inspect

`export_model` takes an `ExportSpec` that states every setting that changes the exported bytes. With a trained `model` and a NumPy array `representative_x` matching its input shape:

```python
from pathlib import Path

import numpy as np
from helia_edge.export import ExportSpec, LiteRTRunner, export_model

spec = ExportSpec(precision="a8w8", io_dtype="int8", mode="concrete")
result = export_model(model, spec, calibration=representative_x.astype(np.float32))
print(result.sha256, result.inputs[0].scale, result.environment.helia_edge)
predictions = LiteRTRunner(result.content).predict(representative_x)
Path("model.tflite").write_bytes(result.content)
```

| Precision | Graph | Input/output dtype |
|---|---|---|
| `fp32` | float32 weights and compute | `float32` |
| `fp16` | native float16 graph | `float16` |
| `a8w8` | int8 activations and weights | `int8` or `float32` |
| `a16w8` | int16 activations, int8 weights | `int16` or `float32` |

- `mode` and `io_dtype` have no defaults, and an invalid precision/dtype pair is rejected.
- `a8w8` and `a16w8` need calibration: a non-empty, finite float32 array shaped like the model input, used one sample at a time in stored order. Other precisions refuse calibration data.
- `litert` is the export format, and it needs the TensorFlow backend. A model with several inputs takes calibration as a mapping of input name to samples and the `keras` or `saved_model` mode (see [Streaming models](https://ambiqai.github.io/helia-edge/guide/export/#streaming-models)). A model trained with the Torch backend is rebuilt from its params and weights in a process started with `KERAS_BACKEND=tensorflow`.
- Choose representative inputs from the distribution the model will encounter. Compare predictions against the original model and evaluate task quality on held-out data before deployment.
- `ExportSpec` values match helia-benchmark's precision names.
- `LiteRTRunner` runs model bytes with ai-edge-litert, rounding and saturating integer inputs. It runs native float16 graphs where the runtime has float16 kernels for every operator.

## Export with a record

`export` exports a model with its export record (`helia-edge/export-record@1`): one JSON file per artifact with the model spec, the weights digest, every export setting, the artifact's sha256 and I/O, and the environment. It names the calibration and the weights by content hash only, never by path. `write` puts `model.tflite`, `model.weights.h5` and `record.json` in a directory, and `load_export_record` rebuilds the Keras model from them, checking the weights against the digest:

```python
from helia_edge.export import export, load_export_record
from helia_edge.models import ModelSpec, build, compact_tcn_params

spec = ModelSpec(params=compact_tcn_params(num_classes=2), input_shape=(256, 8))
model = build(spec)  # trained with a dynamic batch
result = export(model, precision="a8w8", io_dtype="int8", calibration=representative_x, spec=spec)
record_path = result.write("exports/tcn-a8w8")
model_again = load_export_record(record_path)
```

- **Static batch:** the exported model has batch `batch_size` (default 1). With `spec`, the export is `build(spec, batch_size=batch_size)` with the model's weights, whatever batch the model was built with, so a model trained with a dynamic batch exports with no `SPACE_TO_BATCH_ND`/`BATCH_TO_SPACE_ND` pairs or dynamic tensors, and the record re-exports the same bytes in another process. Without `spec`, the model is exported as it is and its batch must be `batch_size`.
- **Weights digest:** sha256 over each weight's dtype, shape and value, in `model.weights` order (not the weight names, which Keras numbers per process for unnamed layers). It is the same whether the weights come from `.weights.h5`, `.keras` or an import (`weights_import` records the mapping and the source file's sha256).
- **Calibrated precisions** (`a8w8`, `a16w8`) export with `batch_size` 1. Streaming `resets` are increasing steps within the calibration.
- **Streaming models:** pass the signal as `calibration` with the state `resets`; `export` builds the state calibration with `stream_calibration`. `with_golden(inputs, resets)` adds a `helia-model-zoo/golden@2` sequence run with LiteRT's reference kernels.
- **Options:** `ExportOptions(strict=True, state_tie_tolerance=0.01, mode="keras", dense_per_channel=True)`; `mode="keras"` keeps the model's batch; `concrete` traces batch 1 and exports `batch_size` 1 only.
- **Dense weights:** for `a8w8` and `a16w8`, `dense_per_channel=False` quantizes `FULLY_CONNECTED` weights with one scale per tensor, while convolutions stay per channel. The CMSIS-NN int16 `FULLY_CONNECTED` kernel of LiteRT for Microcontrollers accepts per-tensor weights only, so `a16w8` exports for it use `False`; the reference and heliaRT kernels run either. The record names the setting, and float precisions refuse `False`.

## Streaming models

A streaming model processes one step per call and carries its recurrent state as explicit inputs and outputs. Name each state pair with `state_input(k, ...)` and `state_output(k, ...)`: the runtime feeds `state_out_k` back as `state_in_k` and feeds zeros at the start of an independent sequence. `StreamingLSTMCell` is one LSTM step with the weight layout of `keras.layers.LSTMCell`.

```python
import keras
import numpy as np
from helia_edge.export import ExportSpec, LiteRTStreamRunner, export_model, stream_calibration
from helia_edge.layers import StreamingLSTMCell, state_input, state_output

keras.utils.set_random_seed(0)
x = keras.Input((16,), batch_size=1, name="features")
h, c = state_input(0, (32,), batch_size=1), state_input(1, (32,), batch_size=1)
h_next, c_next = StreamingLSTMCell(32)([x, h, c])
prob = keras.layers.Dense(1, activation="sigmoid", name="prob")(h_next)
model = keras.Model([x, h, c], [prob, state_output(0, h_next), state_output(1, c_next)])

features = np.random.default_rng(0).normal(size=(200, 16)).astype(np.float32)
calibration = stream_calibration(model, {"features": features})
result = export_model(model, ExportSpec(precision="a16w8", io_dtype="int16", mode="keras"), calibration)
runner = LiteRTStreamRunner(result.content)
fed, outputs = runner.run({"features": runner.encode("features", features[:50, None])})
probabilities = runner.decode("prob", outputs["prob"])
```

- **Mode:** use `keras`. `concrete` converts single-input models only, because its graph has no signature to name inputs and states. `saved_model` exports a dynamic batch, so integer state pairs cannot be tied.
- **Records:** `result.inputs` and `result.outputs` list tensors in the model's subgraph order. State tensors have role `state` and the index `pair` of their pair; the converter may reorder inputs, so address them by name.
- **Calibration:** `stream_calibration` runs the Keras model over the signal inputs and returns every input, the state inputs filled with the states the model produces (zero at the start and at the steps in `resets`).
- **Tied state:** with integer I/O (`a8w8` with `int8`, `a16w8` with `int16`), the converter calibrates `state_in_k` and `state_out_k` separately. Export then gives each pair one scale and zero point whose range covers both tensors' ranges (for `int16`, the larger scale), so a raw `state_out_k` keeps its value as the next `state_in_k`; the bias of a fully connected or convolution layer that reads a state is requantized with it. It refuses when the tied scale differs from either original by more than `ExportSpec.state_tie_tolerance` of it (default 1%, at most 50%), or when an operator that needs the state tensor's parameters unchanged (such as a concatenation, or a tanh writing the state) reads or writes it; a reshape passes the tied parameters on to its output. State inputs and outputs must pair up with equal shapes and types, and every output needs its own name (two outputs of one layer share it). The export record's `io` records `state_scales_tied`.
- **Running:** `LiteRTStreamRunner` runs a stream on raw tensors, carrying the state unchanged; it returns every input as fed and every output, stacked by step. `LiteRTRunner` stays single-input.
- `export` takes streaming models with one signal input; see [Export with a record](https://ambiqai.github.io/helia-edge/guide/export/#export-with-a-record).

## The command line

`helia-edge export create` builds a model from a `ModelSpec` file (YAML or JSON), loads its weights, exports it with `export` and writes the record directory. `helia-edge export reproduce` exports a record's model again and compares:

```sh
KERAS_BACKEND=tensorflow helia-edge export create spec.yaml --weights trained.weights.h5 \
  --precision a8w8 --calibration calibration.npy --out exports/tcn-a8w8
KERAS_BACKEND=tensorflow helia-edge export reproduce exports/tcn-a8w8/record.json \
  --weights exports/tcn-a8w8/model.weights.h5 --calibration calibration.npy
```

- **`create`:** `--weights` is a Keras weights file (`.weights.h5` or `.keras`), or with `--mapping NAME` the source file imported through that mapping of the spec's family (see [Import weights](https://ambiqai.github.io/helia-edge/guide/import-weights/)). `--io-dtype` defaults to `float32`, `float16`, `int8` or `int16` for `fp32`, `fp16`, `a8w8` and `a16w8`. `--calibration` and `--golden-inputs` are `.npy` files, with `--resets` and `--golden-resets` repeated for each reset step. `--batch-size` (default 1), `--mode keras|concrete`, `--no-strict`, `--state-tie-tolerance` and `--dense-per-tensor` are passed to `export`. A missing argument file or a `--batch-size` outside 1 to 2³¹−1 is a usage error (exit 2). A refused export (an invalid spec, setting or input file, a backend other than TensorFlow) exits 1 with an error and writes no record.
- **`reproduce`:** `--weights` is the record's weights (any file with the recorded digest), or for imported weights the source file with the recorded sha256. `--calibration` is required when the record has calibration and refused when it has none; `--golden-inputs` regenerates and compares the golden (without it, the output says the golden was not compared). A `model.tflite` or `golden.npz` beside the record is compared with the record's size and sha256. Exit codes:
  - 0: the weights digest, the export settings, the artifact's sha256 and size, the I/O, the golden (with `--golden-inputs`) and the files beside the record match;
  - 1: something differs, and each difference is listed;
  - 2: the environment differs from the recorded one, and nothing is exported. `--allow-env-mismatch` compares anyway and lists the differences. Also when this process cannot export: Keras is missing or its backend is not TensorFlow;
  - 3: the record is missing, invalid or has no model spec, or an input file is missing, unreadable, does not match its recorded hash or is not used by the record;
  - 4: the export itself failed with an error, printed with its traceback; nothing was compared.
- **`schema`:** prints the record's JSON Schema.
- **Environment:** the record names the helia-edge version, install source and commit, Python, the platform and the NumPy, Keras, TensorFlow and ai-edge-litert versions. Other versions may export other bytes.
- **Provenance:** `environment.helia_edge.source` says whether the version identifies the code that exported:
  - `release`: installed from a package index (a wheel built from a modified tree and served from a private index looks the same, so prefer `vcs` for published exports);
  - `vcs`: installed from a git URL, at the recorded commit (from a local `git+file://` repository, that commit may never have been pushed);
  - `local`: installed from a directory or an archive (local or by URL), or with only egg-info metadata (a source tree, some system packages), none of which the version identifies;
  - `unknown`.

  `create` warns unless it is `release` or `vcs`; with `--require-provenance` it refuses. For exports you publish, install from git at a commit: `uv pip install 'helia-edge @ git+https://github.com/AmbiqAI/helia-edge@<commit>'`.
- **Layer names:** unnamed layers' names become tensor names in the exported bytes, and Keras numbers them per process. `export` with `spec`, `load_export_record` and both commands build with the numbering started from zero, through a private Keras name table. If a Keras release no longer numbers names through it, they warn (`RuntimeWarning`) that the bytes may depend on layers built earlier in the process.

## Hand off to an Ambiq deployment

The exported model is one part of the deployment. Select the inference runtime and check its operator and precision support before integrating it into firmware. Training backend support does not imply that every model can run on every target.

| Keep with the model | Why the deployment needs it |
|---|---|
| Input shape, layout and dtype | Defines how firmware supplies each inference window. |
| Preprocessing parameters | Keeps scaling, filtering and feature extraction consistent with training. |
| Output names and interpretation | Defines class order, thresholds or reconstruction units. |
| Calibration and quantization settings | Explains how the converted artifact was produced. |
| Representative inputs and expected outputs | Enables comparison between host inference and the deployed runtime. |
| Model and source revisions | Identifies the exact artifact being evaluated. |

Validate task quality after conversion, then measure latency and memory with the model on your intended Ambiq board. Host conversion and operation counts do not establish device performance. For streaming models, include the state shapes, initialization and update order described in the model's guide.
