Skip to content
compressionKIT
User guide
HELIA

Deployment Guide

This guide covers the deployment package emitted by compressionkit.export.deploy.export_for_deployment(), the lightweight runtime APIs, HuggingFace loading, and the optional two-stage entropy-coding path.

For the release-grade package requirements that span AI, DSP, HuggingFace publication, and docs generation, see the v1 release contract.

Use load_codec() to select the runtime from package metadata. Start with a downloaded package and a representative frame before splitting the pipeline across devices. See the quickstart for installation and the common interface.

StepYour integration ownsPackage/runtime supplies
Prepare a frameSensor acquisition, sample-rate conversion, frame boundaries and unitsExpected sample rate, frame size and codec configuration
EncodeTarget runtime integration and buffer managementPython reference codec; exported artifacts according to the manifest
Store or transmitPacket format, codec identity/version, sequencing and error handlingFamily-specific payload and decoding side information
ReconstructMatching package selection and handling missing framesDecoder for the matching encoded representation
ValidateApplication acceptance criteria and target measurementsSaved reference data and evaluation tools

compress() returns an EncodedFrame with payload, nbits and side. Preserve the information needed by decompress():

  • RVQ: the payload is a token-index array. Its NumPy storage size is not its packed token bitrate. An integration must define how indices are packed and restore their shape.
  • SPIHT: the payload is a coded bitstream with decoding side information.
  • Hybrid: the SPIHT payload represents the preprocessed signal; use the matching hybrid package and metadata.

This Python object is not a network protocol. Define a stable envelope with the codec/package version, frame identity, payload length and required side information. Count that envelope and transport overhead when estimating bandwidth savings.

What an embedded export does not establish

Section titled “What an embedded export does not establish”

A model file or C header is an integration input. It does not establish a complete firmware codec, a working sensor-to-radio path, or measured latency and memory use on your board. Check the actual package contents, implement any missing codec/transport pieces, and compare target results against the reference runtime before treating the integration as validated.

The standard RVQ deployment flow has three pieces:

  1. Train a model and export the deploy/ directory under results/<run_name>/deploy/.
  2. Integrate the encoder and RVQ codebook into your runtime to turn frames into token indices. Match the exported tensor types and quantization settings.
  3. Decode either locally, on a host, or in the cloud depending on your product architecture.

Use load_codec() for the common interface across families. The lower-level examples below are specific to RVQ.

export_for_deployment() writes a manifest plus the artifacts needed for single-stage inference. Depending on export options, some files are always present and some are optional.

FileRequiredPurposeTypical target
deploy_manifest.jsonYesDeclares file names, tensor shapes, quantization mode, and codebook metadata.Runtime metadata
encoder.tfliteYesINT8 LiteRT encoder used to produce continuous latents from input frames.MCU / edge device
encoder.hYesC header for the quantized encoder blob.MCU firmware
encoder_float32.tfliteYesFloat32 LiteRT encoder for browser and host runtimes without Keras.Browser / x86 / ARM Linux
encoder.kerasYesReference encoder kept in Keras format.Server / offline tools
decoder.kerasYesReference decoder kept in Keras format.Server / offline tools
decoder_float32.tfliteOptional, exported by defaultFloat32 LiteRT decoder for host-side reconstruction without Keras.x86 / ARM Linux
decoder.tfliteOptionalINT8 LiteRT decoder for full on-device reconstruction.MCU / edge device
decoder.hOptionalC header for the INT8 decoder blob.MCU firmware
codebook.npzYesNumPy archive containing RVQ codebook tables.Python runtime
codebook.hYesC header containing RVQ codebook tables.MCU firmware
sample_data.npzOptionalSynthetic or evaluation sample inputs / targets / reconstructions.Validation / demos
demo_recordings.npzReleaseTen real, quality-gated continuous recordings, resampled to the model rate.Browser demos
demo_recordings_manifest.jsonReleaseDataset attribution, license, source offsets, and quality metrics for the demo recordings.Browser demos / compliance
model_card.jsonOptionalMetadata used when publishing to HuggingFace.Release tooling

The manifest is the source of truth. The runtime reads it first, then resolves the encoder, optional decoder, and codebook files from the names listed there.

When exporting vq.get_weights(), declare rvq_num_levels, rvq_use_ema, and rvq_kmeans_init from the trained model’s configuration. Plain RVQ has one codebook matrix per level. EMA RVQ also stores count vectors and embedding-sum matrices, plus an optional k-means flag; these are training state, not additional codebooks. Export rejects weights that do not match the declared layout. The standalone NPZ/header writers accept codebook matrices only; first call extract_codebooks(weights, num_levels=..., use_ema=..., kmeans_init=...) for EMA.

Saved-run repackaging uses the run’s config.json and loads positional weight arrays in numeric order. Preserve original packages and use golden repackage <id> --output-dir <separate-directory> when repairing an existing export.

With sample inputs, reference_vectors.npz contains both deployed-path vectors and independent source_latents, source_indices, source_quantized_latents, and source_reconstructions from the restored training quantizer. Validation requires identical indices and quantized latents for identical source latents; the float encoder and decoder comparisons use atol=rtol=1e-5. This isolates codebook correctness from encoder conversion and precision differences. Sample-data reconstructions also use the discrete trained path rather than bypassing VQ. Training references run on CPU with full float32 operations, so GPU TF32 rounding cannot change the reference or fail the float32 conversion gate. For INT8-only decoder packages, the training comparison uses the exported float Keras decoder companion; the selected INT8 decoder is checked against its own deployed reference vectors. Float tolerances are not applied to INT8 outputs.

Strict release validation rejects older RVQ references without these source arrays. Re-export them from the original training state; a self-consistent deployed reference alone does not establish training parity. Repackaging also recomputes model-card precision measurements instead of copying historical reports that may have used different codebooks.

Correcting the exported level count changes packet interpretation. Keep existing packets paired with their original bundle and bind new packets to the corrected bundle revision; replacing tables inside an existing package is not a migration.

Both RVQ encoders accept one normalized frame: PPG is (1, 1, 320, 1) at 64 Hz and ECG is (1, 1, 512, 1) at 256 Hz. Resample and frame the raw single-channel recording first, then normalize each frame independently:

mean = frame.mean()
scale = np.sqrt(np.mean((frame - mean) ** 2) + 1e-3)
normalized = (frame - mean) / scale

Retain mean and scale to return a decoded frame to its original units: raw = decoded * scale + mean. Do not feed raw ADC or continuous values directly to encoder_int8.tflite; its calibration domain is this training-time normalized representation.

Every INT8 RVQ export reservoir-samples up to 65,536 real preprocessed validation frames, using 4,096 frames for LiteRT calibration and a disjoint 2,048-frame holdout for quantization_report.json. The release gate requires at most 1% input saturation on that holdout. A P90 reconstruction PRD target of 10% against the float32 encoder remains a quality recommendation. Worst-frame PRD remains a 15% tail-risk warning in the report, rather than a release blocker, because isolated RVQ decision-boundary crossings can be disproportionate. compressionkit golden validate-all --strict-release and the Hugging Face publisher reject INT8 RVQ releases without a passed report.

RVQ releases carry ten 30-second real recordings in demo_recordings.npz: BIDMC PPG at 64 Hz and MIT-BIH ECG at 256 Hz. The companion manifest retains ODC-By attribution and the quality gate results. Signals are continuous source waveforms after resampling; apply the model’s normal framing and normalization before inference.

Refresh local bundles before a manual Hugging Face update:

Terminal window
scripts/devcontainer.sh exec -- uv run python scripts/attach_rvq_demo_recordings.py --modality all

To update only validated release candidates, repeat --experiment-id, for example --experiment-id ecg-rvq-2x --experiment-id ecg-rvq-4x.

Use --duration-seconds 120 for two-minute clips. The command updates only existing results/*/deploy/ packages, their manifests, and checksums; it does not train or re-export models.

Use the source-checkout installation from Getting started. LiteRT executes the exported models, but importing the package also loads modules with Keras and other development dependencies; this is not a separately packaged minimal runtime.

import numpy as np
from compressionkit.runtime import RVQCodec
def synthetic_ppg_frame(frame_size: int = 320, sample_rate: int = 64) -> np.ndarray:
t = np.arange(frame_size, dtype=np.float32) / np.float32(sample_rate)
waveform = (
0.55 * np.sin(2.0 * np.pi * 1.2 * t)
+ 0.18 * np.sin(2.0 * np.pi * 2.4 * t + 0.3)
+ 0.03 * np.sin(2.0 * np.pi * 0.15 * t)
)
return waveform.reshape(1, 1, frame_size, 1).astype(np.float32)
codec = RVQCodec("results/ppg_rvq_64hz_04x_golden/deploy")
signal = synthetic_ppg_frame()
indices = codec.encode(signal)
reconstruction = codec.decode(indices)
print(indices.shape)
print(reconstruction.shape)
print(codec.num_levels, codec.num_embeddings, codec.embedding_dim)

Use the lower-level methods when you want to separate the pipeline into explicit steps:

latent = codec.encode_latent(signal)
indices = codec.quantize_latent(latent)
quantized_latent = codec.dequantize_indices(indices)
if codec.has_decoder:
reconstruction = codec.decode_latent(quantized_latent)

This is useful when the encoder runs on-device and the decoder runs elsewhere.

Published model repos follow the convention Ambiq/compressionkit-{modality}-{cr}x-{version}, for example Ambiq/compressionkit-ppg-4x-v1.1 or Ambiq/compressionkit-ecg-8x-v1.1.

Install the optional Hub dependency:

Terminal window
uv sync --extra hf

Then load a codec directly from the Hub:

import numpy as np
from compressionkit.runtime import RVQCodec
def synthetic_ecg_frame(frame_size: int = 512, sample_rate: int = 256) -> np.ndarray:
t = np.arange(frame_size, dtype=np.float32) / np.float32(sample_rate)
waveform = (
0.75 * np.sin(2.0 * np.pi * 1.1 * t)
+ 0.12 * np.sin(2.0 * np.pi * 9.0 * t)
+ 0.02 * np.sin(2.0 * np.pi * 0.3 * t)
)
return waveform.reshape(1, 1, frame_size, 1).astype(np.float32)
codec = RVQCodec.from_pretrained("Ambiq/compressionkit-ecg-4x-v1.1")
signal = synthetic_ecg_frame()
indices = codec.encode(signal)
reconstruction = codec.decode(indices)

The Hub staging step renames a few artifacts for distribution, but RVQCodec.from_pretrained() handles those differences for you. In particular, it maps:

  • config.json to deploy_manifest.json
  • encoder_int8.tflite to encoder.tflite
  • decoder_int8.tflite to decoder.tflite
  • sample_stimulus.npz to sample_data.npz

The two-stage path is for advanced users who want additional bitrate reduction beyond uniform RVQ token coding.

Stage 1 uses the standard RVQ codec to produce token indices. Stage 2 runs a causal entropy prior over those tokens and arithmetic-codes them into a compressed bitstream.

import numpy as np
from compressionkit.runtime import RVQCodec
from compressionkit.runtime.prior import EntropyPrior
from compressionkit.runtime.two_stage import TwoStageCodec
def synthetic_ppg_frame(frame_size: int = 320, sample_rate: int = 64) -> np.ndarray:
t = np.arange(frame_size, dtype=np.float32) / np.float32(sample_rate)
waveform = 0.6 * np.sin(2.0 * np.pi * 1.25 * t) + 0.1 * np.sin(2.0 * np.pi * 2.5 * t)
return waveform.reshape(1, 1, frame_size, 1).astype(np.float32)
run_dir = "results/ppg_rvq_64hz_04x_golden"
codec = RVQCodec(f"{run_dir}/deploy")
prior = EntropyPrior(f"{run_dir}/deploy/prior_int8.tflite")
two_stage = TwoStageCodec(codec, prior)
signal = synthetic_ppg_frame()
result = two_stage.compress(signal)
reconstruction = two_stage.decompress(result)
print(result.bits_per_token_uniform)
print(result.bits_per_token_actual)
print(result.cr_uplift)

If you already have RVQ indices, you can skip the encoder and operate directly on tokens:

indices = codec.encode(signal)
estimate = two_stage.estimate_bitrate(indices)
compressed = two_stage.compress_indices(indices)
restored_indices = two_stage.decompress_indices(compressed)

Use the two-stage path when transport or storage cost is the limiting factor and you can afford the extra prior model. Use the single-stage codec when simplicity, fixed compute, or embedded deployment dominates.

SPIHT and hybrid packages use family-specific runtime loaders. Use load_codec to dispatch from package metadata, and consult the experiment page for the matching bundle target.

  • Keep encoder.tflite and codebook.h on-device.
  • Add decoder.tflite and decoder.h only if the product must reconstruct on the same device.
  • Prefer fixed frame shapes that match the exported encoder input exactly.
  • Validate end-to-end quantization behavior with sample_data.npz before integrating with firmware.
  • LiteRT can execute the exported encoder and decoder; install the package dependencies required by the Python imports.
  • decoder_float32.tflite supports host reconstruction through a LiteRT interpreter rather than loading the Keras decoder model. This does not remove Keras from the package import dependencies.
  • HuggingFace loading is most convenient in notebooks, evaluation services, and server-side decode pipelines.
  • A common production split is encoder plus codebook on the wearable and decoder off-device.
  • In that arrangement, transmit RVQ indices or the two-stage bitstream instead of raw windows.
  • The runtime helpers (encode_latent, quantize_latent, dequantize_indices, decode_latent) map cleanly onto that boundary.

Before shipping a deploy package, verify these basics:

  1. Load the deploy/ directory with RVQCodec and confirm codec.manifest matches the expected frame size and latent shape.
  2. Run encode() and decode() on a synthetic frame to confirm artifact completeness.
  3. If you publish to HuggingFace, verify RVQCodec.from_pretrained() on a clean environment.
  4. If you ship a two-stage path, validate both compress() and decompress() with the exact prior artifact you plan to release.