Skip to content
compressionKIT
User guide
HELIA

RVQ Autoencoder

The Residual Vector Quantization (RVQ) Autoencoder is the primary compression method in compressionKIT. It combines a convolutional encoder/decoder with a multi-level discrete bottleneck to achieve high compression ratios while maintaining signal fidelity.

The encoder uses a series of stride-2 stages to downsample the input temporally:

  • First 2 stages: Standard Conv2D blocks (kernel 7, stride 2)
  • Remaining stages: Depthwise-separable Conv2D blocks (more efficient)
  • Head projection: 1×1 convolution to embedding_dim channels

Each block includes configurable normalization (batch, layer, or none) and ReLU activation.

The total downsampling factor is 2num_stages2^{\text{num\_stages}}.

The Residual Vector Quantizer from heliaEDGE discretizes the continuous latent representation:

  1. Find nearest codebook entry for each latent position
  2. Compute residual (what the first codebook missed)
  3. Quantize the residual with the next codebook
  4. Repeat for MM levels

Each level uses a codebook of size KK (the latent_width parameter). Training uses the straight-through estimator for gradient flow, with commitment and codebook losses.

The decoder mirrors the encoder with upsampling stages:

  • UpSampling2D (2×) → Conv2D → SeparableConv2D (anti-aliasing)
  • Optional normalization per block
  • Final 1×1 convolution to output channels
model:
embedding_dim: 16 # Latent channel dimension
latent_width: 512 # Codebook size K
num_levels: 2 # RVQ levels M
num_stages: 3 # Encoder stages (2^3 = 8× downsample)
base_filters: 48 # First stage filter count
multiplier: 1.25 # Filter growth per stage
beta: 0.25 # VQ commitment loss weight
encoder_block_norm: batch
encoder_head_norm: none
decoder_block_norm: none
decoder_head_norm: layer

The compression ratio depends on the input bit depth, downsampling factor, codebook size, and number of levels:

CR=T×BT2N×M×log⁡2(K)\text{CR} = \frac{T \times B}{\frac{T}{2^N} \times M \times \log_2(K)}
ParameterSymbolPPG ValueECG Value
Frame sizeTT320512
Input bit depthBB1616
Num stagesNN1–41–4
Num levelsMM1–21–2
Codebook sizeKK256256

Example (PPG 8×): CR=320×1640×2×8=8×\text{CR} = \frac{320 \times 16}{40 \times 2 \times 8} = 8\times

PPG and ECG share the following RVQ configurations. ECG also has a 64× configuration; see the ECG model page.

NameStagesLevelsCodebookCRUse Case
02x122562×Maximum fidelity
04x222564×High quality
08x322568×Compare against your quality requirements
16x4225616×Bandwidth-constrained
32x4125632×Extreme compression

See the Model Zoo for full results.

The training pipeline:

  1. Loads signal data (PPG or ECG) via TFRecord cache
  2. Applies preprocessing (random crop + layer norm) and augmentation (Gaussian noise)
  3. Builds the RVQ autoencoder using heliaEDGE components
  4. Trains with Adam optimizer, MSE + derivative loss + RVQ commitment/codebook losses
  5. Monitors val_mse for checkpointing and early stopping
  6. Exports best encoder to INT8 TFLite + C header

The trained encoder is exported as:

  • encoder.tflite — INT8 quantized TFLite model for on-device inference
  • encoder.h — C header with the model weights as a byte array
  • encoder_float32.tflite — FP32 LiteRT encoder for browser and host integrations

The decoder and RVQ codebooks are stored separately for server-side reconstruction. A split deployment runs the encoder and codebook lookup on the device and transmits indices to a separate decoder.