# helia_edge.models.silero_vad

Silero VAD v6 (16 kHz) as a streaming Keras model with explicit LSTM state.

Follows snakers4/silero-vad v6.2.2 (commit 60b7ffa), ``silero_vad_16k_op15.onnx``. One call takes
576 samples, the last 64 samples of the previous call followed by 512 new ones (32 ms at 16 kHz), and
returns the speech probability. The LSTM state is carried as ``state_in_0``/``state_in_1`` (h, c),
fed from the previous call's ``state_out_0``/``state_out_1``; zeros (and zero context) start an
independent stream. Weights come from ``helia_edge.importers`` with ``SILERO_VAD_V6_ONNX``
(``silero_vad_params``).

The default options compute the reference model exactly: an STFT magnitude with a square root. The
other options (``SileroVadParams``) compute the same model in other ways: ``max_projection`` needs no
square root, so the model exports to int16 activations; ``conv_blocks`` and ``live_taps`` are exact
rewrites. They derive what they need (block kernels, the folded padding, tap slices) from the same
weights inside the graph, so an exporter folds them into constants.

## helia_edge.models.silero_vad.SAMPLES

`constant` · `python`

```python
SAMPLES = 576
```

Samples per call: 64 of context and 512 new ones.

Source: `helia_edge/models/silero_vad.py:27`

## helia_edge.models.silero_vad.UNITS

`constant` · `python`

```python
UNITS = 128
```

Source: `helia_edge/models/silero_vad.py:29`

## helia_edge.models.silero_vad.SileroBlockStft

`class` · `python`

```python
SileroBlockStft(magnitude: str = 'sqrt', **kwargs={})
```

STFT magnitude of one call (576 samples) from convolutions over 64-sample blocks.

The same transform as ``StftMagnitude(256, 128, 129, padding=(0, 64))`` with the same stored basis
(256, 1, 258): frames 0 to 2 are a convolution over blocks with stride 2, and the last frame's right
reflect padding is folded into its kernel, so no reflect pad or strided frames over samples remain.
The kernels are derived from the basis in the graph and an exporter folds them into constants. The
output is (batch, 4, 1, 129), the frames as rows of an image, which the encoder convolves without
reshapes.

**Parameters**

| Name | Type | Default | Description |
| --- | --- | --- | --- |
| magnitude | str | 'sqrt' | ``sqrt`` (exact) or ``max_projection`` (no square root, within 0.25%). |

Source: `helia_edge/models/silero_vad.py:33`

### helia_edge.models.silero_vad.SileroBlockStft.magnitude

`attribute` · `python`

```python
magnitude = magnitude
```

Source: `helia_edge/models/silero_vad.py:52`

### helia_edge.models.silero_vad.SileroBlockStft.build

`method` · `python`

```python
build(input_shape)
```

Source: `helia_edge/models/silero_vad.py:54`

### helia_edge.models.silero_vad.SileroBlockStft.call

`method` · `python`

```python
call(audio)
```

Source: `helia_edge/models/silero_vad.py:59`

### helia_edge.models.silero_vad.SileroBlockStft.compute_output_shape

`method` · `python`

```python
compute_output_shape(input_shape)
```

Source: `helia_edge/models/silero_vad.py:81`

### helia_edge.models.silero_vad.SileroBlockStft.get_config

`method` · `python`

```python
get_config()
```

Source: `helia_edge/models/silero_vad.py:84`

## helia_edge.models.silero_vad.SileroFrameConv

`class` · `python`

```python
SileroFrameConv(filters: int, strides: int = 1, **kwargs={})
```

An encoder ``Conv1D`` with ReLU over frames held as rows of a (frames, 1, channels) image.

The kernel keeps the ``Conv1D`` shape (3, channels, filters); input padding is a separate layer.

Source: `helia_edge/models/silero_vad.py:88`

### helia_edge.models.silero_vad.SileroFrameConv.filters

`attribute` · `python`

```python
filters = filters
```

Source: `helia_edge/models/silero_vad.py:99`

### helia_edge.models.silero_vad.SileroFrameConv.strides

`attribute` · `python`

```python
strides = strides
```

Source: `helia_edge/models/silero_vad.py:100`

### helia_edge.models.silero_vad.SileroFrameConv.build

`method` · `python`

```python
build(input_shape)
```

Source: `helia_edge/models/silero_vad.py:102`

### helia_edge.models.silero_vad.SileroFrameConv.call

`method` · `python`

```python
call(x)
```

Source: `helia_edge/models/silero_vad.py:108`

### helia_edge.models.silero_vad.SileroFrameConv.compute_output_shape

`method` · `python`

```python
compute_output_shape(input_shape)
```

Source: `helia_edge/models/silero_vad.py:112`

### helia_edge.models.silero_vad.SileroFrameConv.get_config

`method` · `python`

```python
get_config()
```

Source: `helia_edge/models/silero_vad.py:116`

## helia_edge.models.silero_vad.SileroLiveTaps

`class` · `python`

```python
SileroLiveTaps(filters: int, taps: tuple[int, ...], **kwargs={})
```

A zero-padded encoder convolution as a dense layer over the kernel taps that see real frames.

The kernel keeps the convolution's shape (3, channels, filters). ``taps`` lists the taps that see
the input's frames, in order; the input is those frames' channels, flattened.

Source: `helia_edge/models/silero_vad.py:120`

### helia_edge.models.silero_vad.SileroLiveTaps.filters

`attribute` · `python`

```python
filters = filters
```

Source: `helia_edge/models/silero_vad.py:132`

### helia_edge.models.silero_vad.SileroLiveTaps.taps

`attribute` · `python`

```python
taps = tuple(taps)
```

Source: `helia_edge/models/silero_vad.py:133`

### helia_edge.models.silero_vad.SileroLiveTaps.build

`method` · `python`

```python
build(input_shape)
```

Source: `helia_edge/models/silero_vad.py:135`

### helia_edge.models.silero_vad.SileroLiveTaps.call

`method` · `python`

```python
call(x)
```

Source: `helia_edge/models/silero_vad.py:142`

### helia_edge.models.silero_vad.SileroLiveTaps.compute_output_shape

`method` · `python`

```python
compute_output_shape(input_shape)
```

Source: `helia_edge/models/silero_vad.py:146`

### helia_edge.models.silero_vad.SileroLiveTaps.get_config

`method` · `python`

```python
get_config()
```

Source: `helia_edge/models/silero_vad.py:149`

## helia_edge.models.silero_vad.build

`function` · `python`

```python
build(
    params: SileroVadParams,
    input_shape: tuple[int | None, ...] | None = None,
    *,
    batch_size: int | None = None,
    name: str | None = None,
) -> keras.Model
```

Build the Silero VAD v6 16 kHz streaming model, untrained.

Inputs are ``audio`` (576,) float32 in [-1, 1] and the state ``state_in_0`` and ``state_in_1``
(128,). Outputs are ``prob`` (1,) and ``state_out_0`` and ``state_out_1``. Streaming and export use
``batch_size=1``.

**Parameters**

| Name | Type | Default | Description |
| --- | --- | --- | --- |
| params | SileroVadParams | Required | Model parameters: the geometry fixed by the v6.2.2 weights, and the options. |
| input_shape | tuple[int \| None, ...] \| None | None | None, or the audio shape ``(576,)``. |
| batch_size | int \| None | None | Static batch size; None for a dynamic batch. |
| name | str \| None | None | Model name; the family when None. |

**Returns**

| Name | Type | Description |
| --- | --- | --- |
|  | keras.Model | keras.Model: The model, named ``silero_vad`` unless ``name`` is given. |

Source: `helia_edge/models/silero_vad.py:153`
