# Import weights

`helia_edge.importers` sets every weight of a Keras model from a file trained elsewhere. A `WeightMapping` lists, for each weight, the source tensors and how they become the Keras layout. `import_weights` checks the whole mapping first and assigns nothing unless every check passes.

## Import a pinned model

Each family's mappings live beside its params, as `MAPPINGS` in `helia_edge.models.<family>_params`, keyed by mapping name. heliaEDGE includes two: Silero VAD v6 from the v6.2.2 ONNX file (`SILERO_VAD_V6_ONNX`) and FastEnhancer-T from the `onnx-vd-v1.0.0` ONNX release (`FASTENHANCER_T_ONNX`, see [FastEnhancer](https://ambiqai.github.io/helia-edge/examples/fastenhancer/)). Download the file yourself; heliaEDGE never downloads sources:

```python
from helia_edge.importers import import_weights
from helia_edge.models import SILERO_VAD_V6_ONNX, ModelSpec, SileroVadParams, build

model = build(ModelSpec(params=SileroVadParams()), batch_size=1)
report = import_weights(model, SILERO_VAD_V6_ONNX, "silero_vad_16k_op15.onnx")
print(report.source_sha256, len(report.assignments))
```

The model is a streaming model (see [Streaming models](https://ambiqai.github.io/helia-edge/guide/export/#streaming-models)). By default it computes the reference exactly and exports to `fp32`; LiteRT cannot quantize its STFT magnitude (a square root) to int16 activations, so `a16w8` with the default `strict: true` is refused. Each call takes 576 samples at 16 kHz, the last 64 samples of the previous call followed by 512 new ones, and carries its LSTM state as `state_in_0`/`state_in_1` and `state_out_0`/`state_out_1`. Reading ONNX files needs the `onnx` extra (`uv sync --extra onnx`).

`SileroVadParams` options compute the same model in other ways. Every option has the same weights (paths and shapes), so `SILERO_VAD_V6_ONNX` imports into each; Keras `.weights.h5` files are keyed by layer class, so save one per option set.

- **`magnitude="max_projection"`:** each bin's magnitude is the largest of 9 projections of (|re|, |im|), within 0.25% in float32. It is sqrt-free, so the model exports to `a16w8` with `int16` I/O. With `magnitude="sqrt"`, the other options still refuse `int16`.
- **`stft="conv_blocks"`:** the STFT runs as convolutions over 64-sample blocks, with the right reflect padding folded into the last frame. It is exact in float and mirror-pad-free. It is also faster on some targets: for example, Vela estimates about a third of the Ethos-U cycles of strided frames, for a larger model.
- **`encoder_tail="live_taps"`:** the last two encoder layers run as dense layers over the kernel taps that see real frames. It is exact in float, and an export drops the weights of taps that only see padding.

The derived kernels are computed from the weights in the graph, and the export folds them into constants. All three together make the model sqrt-free and mirror-pad-free: its export has no SQRT, MIRROR_PAD or TRANSPOSE operators. Use them for toolchains without an int16 square root (such as Vela for Ethos-U) or without reflect padding: `SileroVadParams(stft="conv_blocks", magnitude="max_projection", encoder_tail="live_taps")`. Calibrate the audio input with full-scale speech so its int16 range covers [-1, 1].

`helia-edge export create` imports the same way, naming the mapping: `helia-edge export create silero.yaml --mapping silero_vad_v6_onnx --weights silero_vad_16k_op15.onnx --precision fp32 --out exports/silero`. The record names the mapping and the source file's sha256, and `export reproduce` takes the same source file (see [The command line](https://ambiqai.github.io/helia-edge/guide/export/#the-command-line)).

To check the import, the repository's parity test compares the model with ONNX Runtime with the state carried: set `HELIA_EDGE_SILERO_ONNX` to the file, and `HELIA_EDGE_SILERO_AUDIO` to 16 kHz speech (a 16-bit mono WAV) to also compare speech decisions, then run `pytest tests/models/test_silero_vad.py`.

## What is checked

- **Source:** the file's sha256 must equal the mapping's pinned `source.sha256`. A file with the same name but other weights is refused: for example, Silero's `silero_vad_16k.safetensors` holds different weights from the v6.2.2 ONNX.
- **Sources used once:** every tensor in the file is used exactly once, or listed in `unused`. A tensor split into parts must have each part used once.
- **Weights set once:** every weight of the model is assigned exactly once.
- **Values:** sources must have their row's `source_shape`, when given, and shapes must match after the transforms, float weights take float sources and must be finite as stored (for example in bfloat16), and integer weights must stay within their range.

## Write a mapping

```python
from helia_edge.importers import GateReorder, Reshape, SourcePin, SumParts, Transpose, WeightMapping, WeightRow

mapping = WeightMapping(
    name="detector_onnx",
    source=SourcePin(uri="https://example.com/detector.onnx", sha256="<sha256>", format="onnx"),
    rows=(
        WeightRow(sources=("conv.weight",), transforms=(Transpose(perm=(2, 1, 0)),), layer="conv", weight="kernel"),
        WeightRow(sources=("conv.bias",), layer="conv", weight="bias"),
        # An ONNX LSTM (hidden 128, input 64): W is (1, 512, 64) and B is (1, 1024), input then recurrent bias
        WeightRow(
            sources=("lstm.W",),
            transforms=(Reshape(shape=(512, 64)), GateReorder(source="iofc", target="ifco"), Transpose(perm=(1, 0))),
            layer="lstm",
            weight="kernel",
        ),
        WeightRow(
            sources=("lstm.B",),
            transforms=(Reshape(shape=(1024,)), SumParts(axis=0, parts=2), GateReorder(source="iofc", target="ifco")),
            layer="lstm",
            weight="bias",
        ),
    ),
)
```

- **Rows:** a row names the Keras layer (`outer/inner` inside a nested model) and the weight within it, such as `kernel` or `bias`, or its path inside a composite layer, such as `query/kernel` in a `MultiHeadAttention`. An optional `source_shape` checks each source tensor's shape before the transforms; give it whenever a row starts with a `Reshape`, which would otherwise accept a transposed source of the same size.
- **Combining sources:** several sources are combined with `sum` or `concat` before the transforms.
- **Transforms:** `Transpose`, `Reshape`, `GateReorder` (for example ONNX LSTM gates `iofc` to Keras `ifco`; PyTorch's `ifgo` is already the Keras order), `SumParts` (add equal parts, such as an ONNX LSTM's two biases) and `Split`, which may only come first.
- **Formats:** the source pin's `format` is one of `SourceFormat`. `onnx` reads the initializers of a self-contained file (not `Constant` nodes; weights in separate data files are refused, since the pin covers one file), `safetensors` reads a `.safetensors` file with NumPy, and `torch` reads a flat state dict saved with `torch.save` (loaded with `weights_only=True`).
- **Serialization:** mappings are pydantic models, so they serialize to and from JSON.
