Import weights
helia_edge.importers sets every weight of a Keras model from a file trained elsewhere. A WeightMapping lists, for each weight, the source tensors and how they become the Keras layout. import_weights checks the whole mapping first and assigns nothing unless every check passes.
Import a pinned model
Section titled “Import a pinned model”Each family’s mappings live beside its params, as MAPPINGS in helia_edge.models.<family>_params, keyed by mapping name. heliaEDGE includes two: Silero VAD v6 from the v6.2.2 ONNX file (SILERO_VAD_V6_ONNX) and FastEnhancer-T from the onnx-vd-v1.0.0 ONNX release (FASTENHANCER_T_ONNX, see FastEnhancer). Download the file yourself; heliaEDGE never downloads sources:
from helia_edge.importers import import_weightsfrom helia_edge.models import SILERO_VAD_V6_ONNX, ModelSpec, SileroVadParams, build
model = build(ModelSpec(params=SileroVadParams()), batch_size=1)report = import_weights(model, SILERO_VAD_V6_ONNX, "silero_vad_16k_op15.onnx")print(report.source_sha256, len(report.assignments))The model is a streaming model (see Streaming models). By default it computes the reference exactly and exports to fp32; LiteRT cannot quantize its STFT magnitude (a square root) to int16 activations, so a16w8 with the default strict: true is refused. Each call takes 576 samples at 16 kHz, the last 64 samples of the previous call followed by 512 new ones, and carries its LSTM state as state_in_0/state_in_1 and state_out_0/state_out_1. Reading ONNX files needs the onnx extra (uv sync --extra onnx).
SileroVadParams options compute the same model in other ways. Every option has the same weights (paths and shapes), so SILERO_VAD_V6_ONNX imports into each; Keras .weights.h5 files are keyed by layer class, so save one per option set.
magnitude="max_projection": each bin’s magnitude is the largest of 9 projections of (|re|, |im|), within 0.25% in float32. It is sqrt-free, so the model exports toa16w8withint16I/O. Withmagnitude="sqrt", the other options still refuseint16.stft="conv_blocks": the STFT runs as convolutions over 64-sample blocks, with the right reflect padding folded into the last frame. It is exact in float and mirror-pad-free. It is also faster on some targets: for example, Vela estimates about a third of the Ethos-U cycles of strided frames, for a larger model.encoder_tail="live_taps": the last two encoder layers run as dense layers over the kernel taps that see real frames. It is exact in float, and an export drops the weights of taps that only see padding.
The derived kernels are computed from the weights in the graph, and the export folds them into constants. All three together make the model sqrt-free and mirror-pad-free: its export has no SQRT, MIRROR_PAD or TRANSPOSE operators. Use them for toolchains without an int16 square root (such as Vela for Ethos-U) or without reflect padding: SileroVadParams(stft="conv_blocks", magnitude="max_projection", encoder_tail="live_taps"). Calibrate the audio input with full-scale speech so its int16 range covers [-1, 1].
helia-edge export create imports the same way, naming the mapping: helia-edge export create silero.yaml --mapping silero_vad_v6_onnx --weights silero_vad_16k_op15.onnx --precision fp32 --out exports/silero. The record names the mapping and the source file’s sha256, and export reproduce takes the same source file (see The command line).
To check the import, the repository’s parity test compares the model with ONNX Runtime with the state carried: set HELIA_EDGE_SILERO_ONNX to the file, and HELIA_EDGE_SILERO_AUDIO to 16 kHz speech (a 16-bit mono WAV) to also compare speech decisions, then run pytest tests/models/test_silero_vad.py.
What is checked
Section titled “What is checked”- Source: the file’s sha256 must equal the mapping’s pinned
source.sha256. A file with the same name but other weights is refused: for example, Silero’ssilero_vad_16k.safetensorsholds different weights from the v6.2.2 ONNX. - Sources used once: every tensor in the file is used exactly once, or listed in
unused. A tensor split into parts must have each part used once. - Weights set once: every weight of the model is assigned exactly once.
- Values: sources must have their row’s
source_shape, when given, and shapes must match after the transforms, float weights take float sources and must be finite as stored (for example in bfloat16), and integer weights must stay within their range.
Write a mapping
Section titled “Write a mapping”from helia_edge.importers import GateReorder, Reshape, SourcePin, SumParts, Transpose, WeightMapping, WeightRow
mapping = WeightMapping( name="detector_onnx", source=SourcePin(uri="https://example.com/detector.onnx", sha256="<sha256>", format="onnx"), rows=( WeightRow(sources=("conv.weight",), transforms=(Transpose(perm=(2, 1, 0)),), layer="conv", weight="kernel"), WeightRow(sources=("conv.bias",), layer="conv", weight="bias"), # An ONNX LSTM (hidden 128, input 64): W is (1, 512, 64) and B is (1, 1024), input then recurrent bias WeightRow( sources=("lstm.W",), transforms=(Reshape(shape=(512, 64)), GateReorder(source="iofc", target="ifco"), Transpose(perm=(1, 0))), layer="lstm", weight="kernel", ), WeightRow( sources=("lstm.B",), transforms=(Reshape(shape=(1024,)), SumParts(axis=0, parts=2), GateReorder(source="iofc", target="ifco")), layer="lstm", weight="bias", ), ),)- Rows: a row names the Keras layer (
outer/innerinside a nested model) and the weight within it, such askernelorbias, or its path inside a composite layer, such asquery/kernelin aMultiHeadAttention. An optionalsource_shapechecks each source tensor’s shape before the transforms; give it whenever a row starts with aReshape, which would otherwise accept a transposed source of the same size. - Combining sources: several sources are combined with
sumorconcatbefore the transforms. - Transforms:
Transpose,Reshape,GateReorder(for example ONNX LSTM gatesiofcto Kerasifco; PyTorch’sifgois already the Keras order),SumParts(add equal parts, such as an ONNX LSTM’s two biases) andSplit, which may only come first. - Formats: the source pin’s
formatis one ofSourceFormat.onnxreads the initializers of a self-contained file (notConstantnodes; weights in separate data files are refused, since the pin covers one file),safetensorsreads a.safetensorsfile with NumPy, andtorchreads a flat state dict saved withtorch.save(loaded withweights_only=True). - Serialization: mappings are pydantic models, so they serialize to and from JSON.