SAMPLES
PythonSamples per call: 64 of context and 512 new ones.
SAMPLES = 576Samples per call: 64 of context and 512 new ones.
Silero VAD v6 (16 kHz) as a streaming Keras model with explicit LSTM state.
Follows snakers4/silero-vad v6.2.2 (commit 60b7ffa), silero_vad_16k_op15.onnx. One call takes
576 samples, the last 64 samples of the previous call followed by 512 new ones (32 ms at 16 kHz), and
returns the speech probability. The LSTM state is carried as state_in_0/state_in_1 (h, c),
fed from the previous call’s state_out_0/state_out_1; zeros (and zero context) start an
independent stream. Weights come from helia_edge.importers with SILERO_VAD_V6_ONNX
(silero_vad_params).
The default options compute the reference model exactly: an STFT magnitude with a square root. The
other options (SileroVadParams) compute the same model in other ways: max_projection needs no
square root, so the model exports to int16 activations; conv_blocks and live_taps are exact
rewrites. They derive what they need (block kernels, the folded padding, tap slices) from the same
weights inside the graph, so an exporter folds them into constants.
SAMPLESconstantSamples per call: 64 of context and 512 new ones.UNITSconstantSileroBlockStftclassSTFT magnitude of one call (576 samples) from convolutions over 64-sample blocks.SileroFrameConvclassAn encoder Conv1D with ReLU over frames held as rows of a (frames, 1, channels) image.SileroLiveTapsclassA zero-padded encoder convolution as a dense layer over the kernel taps that see real frames.buildfunctionBuild the Silero VAD v6 16 kHz streaming model, untrained.Samples per call: 64 of context and 512 new ones.
SAMPLES = 576Samples per call: 64 of context and 512 new ones.
UNITS = 128STFT magnitude of one call (576 samples) from convolutions over 64-sample blocks.
SileroBlockStft(magnitude: str = 'sqrt', **kwargs={})STFT magnitude of one call (576 samples) from convolutions over 64-sample blocks.
The same transform as StftMagnitude(256, 128, 129, padding=(0, 64)) with the same stored basis
(256, 1, 258): frames 0 to 2 are a convolution over blocks with stride 2, and the last frame’s right
reflect padding is folded into its kernel, so no reflect pad or strided frames over samples remain.
The kernels are derived from the basis in the graph and an exporter folds them into constants. The
output is (batch, 4, 1, 129), the frames as rows of an image, which the encoder convolves without
reshapes.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
magnitude | str | 'sqrt' | ``sqrt`` (exact) or ``max_projection`` (no square root, within 0.25%). |
magnitude = magnitudebuild(input_shape)call(audio)compute_output_shape(input_shape)get_config()An encoder Conv1D with ReLU over frames held as rows of a (frames, 1, channels) image.
SileroFrameConv(filters: int, strides: int = 1, **kwargs={})An encoder Conv1D with ReLU over frames held as rows of a (frames, 1, channels) image.
The kernel keeps the Conv1D shape (3, channels, filters); input padding is a separate layer.
filters = filtersstrides = stridesbuild(input_shape)call(x)compute_output_shape(input_shape)get_config()A zero-padded encoder convolution as a dense layer over the kernel taps that see real frames.
SileroLiveTaps(filters: int, taps: tuple[int, ...], **kwargs={})A zero-padded encoder convolution as a dense layer over the kernel taps that see real frames.
The kernel keeps the convolution’s shape (3, channels, filters). taps lists the taps that see
the input’s frames, in order; the input is those frames’ channels, flattened.
filters = filterstaps = tuple(taps)build(input_shape)call(x)compute_output_shape(input_shape)get_config()Build the Silero VAD v6 16 kHz streaming model, untrained.
build( params: SileroVadParams, input_shape: tuple[int | None, ...] | None = None, *, batch_size: int | None = None, name: str | None = None,) -> keras.ModelBuild the Silero VAD v6 16 kHz streaming model, untrained.
Inputs are audio (576,) float32 in [-1, 1] and the state state_in_0 and state_in_1
(128,). Outputs are prob (1,) and state_out_0 and state_out_1. Streaming and export use
batch_size=1.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
params | SileroVadParams | Required | Model parameters: the geometry fixed by the v6.2.2 weights, and the options. |
input_shape | tuple[int | None, ...] | None | None | None, or the audio shape ``(576,)``. |
batch_size | int | None | None | Static batch size; None for a dynamic batch. |
name | str | None | None | Model name; the family when None. |
Returns
| Type | Description |
|---|---|
keras.Model | keras.Model: The model, named ``silero_vad`` unless ``name`` is given. |