Deployment Guide
This guide covers the deployment package emitted by compressionkit.export.deploy.export_for_deployment(), the lightweight runtime APIs, HuggingFace loading, and the optional two-stage entropy-coding path.
For the release-grade package requirements that span AI, DSP, HuggingFace publication, and docs generation, see the v1 release contract.
Start with a host round-trip
Section titled “Start with a host round-trip”Use load_codec() to select the runtime from package metadata. Start with a downloaded package and a representative frame before splitting the pipeline across devices. See the quickstart for installation and the common interface.
From sensor to reconstruction
Section titled “From sensor to reconstruction”| Step | Your integration owns | Package/runtime supplies |
|---|---|---|
| Prepare a frame | Sensor acquisition, sample-rate conversion, frame boundaries and units | Expected sample rate, frame size and codec configuration |
| Encode | Target runtime integration and buffer management | Python reference codec; exported artifacts according to the manifest |
| Store or transmit | Packet format, codec identity/version, sequencing and error handling | Family-specific payload and decoding side information |
| Reconstruct | Matching package selection and handling missing frames | Decoder for the matching encoded representation |
| Validate | Application acceptance criteria and target measurements | Saved reference data and evaluation tools |
What crosses the transport boundary
Section titled “What crosses the transport boundary”compress() returns an EncodedFrame with payload, nbits and side. Preserve the information needed by decompress():
- RVQ: the payload is a token-index array. Its NumPy storage size is not its packed token bitrate. An integration must define how indices are packed and restore their shape.
- SPIHT: the payload is a coded bitstream with decoding side information.
- Hybrid: the SPIHT payload represents the preprocessed signal; use the matching hybrid package and metadata.
This Python object is not a network protocol. Define a stable envelope with the codec/package version, frame identity, payload length and required side information. Count that envelope and transport overhead when estimating bandwidth savings.
What an embedded export does not establish
Section titled “What an embedded export does not establish”A model file or C header is an integration input. It does not establish a complete firmware codec, a working sensor-to-radio path, or measured latency and memory use on your board. Check the actual package contents, implement any missing codec/transport pieces, and compare target results against the reference runtime before treating the integration as validated.
Deployment Workflow
Section titled “Deployment Workflow”The standard RVQ deployment flow has three pieces:
- Train a model and export the
deploy/directory underresults/<run_name>/deploy/. - Integrate the encoder and RVQ codebook into your runtime to turn frames into token indices. Match the exported tensor types and quantization settings.
- Decode either locally, on a host, or in the cloud depending on your product architecture.
Use load_codec() for the common interface across families. The lower-level examples below are specific to RVQ.
Deploy Package Contents
Section titled “Deploy Package Contents”export_for_deployment() writes a manifest plus the artifacts needed for single-stage inference. Depending on export options, some files are always present and some are optional.
| File | Required | Purpose | Typical target |
|---|---|---|---|
deploy_manifest.json | Yes | Declares file names, tensor shapes, quantization mode, and codebook metadata. | Runtime metadata |
encoder.tflite | Yes | INT8 LiteRT encoder used to produce continuous latents from input frames. | MCU / edge device |
encoder.h | Yes | C header for the quantized encoder blob. | MCU firmware |
encoder_float32.tflite | Yes | Float32 LiteRT encoder for browser and host runtimes without Keras. | Browser / x86 / ARM Linux |
encoder.keras | Yes | Reference encoder kept in Keras format. | Server / offline tools |
decoder.keras | Yes | Reference decoder kept in Keras format. | Server / offline tools |
decoder_float32.tflite | Optional, exported by default | Float32 LiteRT decoder for host-side reconstruction without Keras. | x86 / ARM Linux |
decoder.tflite | Optional | INT8 LiteRT decoder for full on-device reconstruction. | MCU / edge device |
decoder.h | Optional | C header for the INT8 decoder blob. | MCU firmware |
codebook.npz | Yes | NumPy archive containing RVQ codebook tables. | Python runtime |
codebook.h | Yes | C header containing RVQ codebook tables. | MCU firmware |
sample_data.npz | Optional | Synthetic or evaluation sample inputs / targets / reconstructions. | Validation / demos |
demo_recordings.npz | Release | Ten real, quality-gated continuous recordings, resampled to the model rate. | Browser demos |
demo_recordings_manifest.json | Release | Dataset attribution, license, source offsets, and quality metrics for the demo recordings. | Browser demos / compliance |
model_card.json | Optional | Metadata used when publishing to HuggingFace. | Release tooling |
The manifest is the source of truth. The runtime reads it first, then resolves the encoder, optional decoder, and codebook files from the names listed there.
RVQ Codebook Export and Training Parity
Section titled “RVQ Codebook Export and Training Parity”When exporting vq.get_weights(), declare rvq_num_levels, rvq_use_ema, and
rvq_kmeans_init from the trained model’s configuration. Plain RVQ has one
codebook matrix per level. EMA RVQ also stores count vectors and embedding-sum
matrices, plus an optional k-means flag; these are training state, not additional
codebooks. Export rejects weights that do not match the declared layout. The
standalone NPZ/header writers accept codebook matrices only; first call
extract_codebooks(weights, num_levels=..., use_ema=..., kmeans_init=...) for EMA.
Saved-run repackaging uses the run’s config.json and loads positional weight
arrays in numeric order. Preserve original packages and use golden repackage <id> --output-dir <separate-directory> when repairing an existing export.
With sample inputs, reference_vectors.npz contains both deployed-path vectors
and independent source_latents, source_indices, source_quantized_latents,
and source_reconstructions from the restored training quantizer. Validation
requires identical indices and quantized latents for identical source latents;
the float encoder and decoder comparisons use atol=rtol=1e-5. This isolates codebook
correctness from encoder conversion and precision differences. Sample-data
reconstructions also use the discrete trained path rather than bypassing VQ.
Training references run on CPU with full float32 operations, so GPU TF32 rounding
cannot change the reference or fail the float32 conversion gate.
For INT8-only decoder packages, the training comparison uses the exported float
Keras decoder companion; the selected INT8 decoder is checked against its own
deployed reference vectors. Float tolerances are not applied to INT8 outputs.
Strict release validation rejects older RVQ references without these source arrays. Re-export them from the original training state; a self-consistent deployed reference alone does not establish training parity. Repackaging also recomputes model-card precision measurements instead of copying historical reports that may have used different codebooks.
Correcting the exported level count changes packet interpretation. Keep existing packets paired with their original bundle and bind new packets to the corrected bundle revision; replacing tables inside an existing package is not a migration.
Encoder Preprocessing and INT8 Parity
Section titled “Encoder Preprocessing and INT8 Parity”Both RVQ encoders accept one normalized frame: PPG is (1, 1, 320, 1) at 64 Hz and ECG is (1, 1, 512, 1) at 256 Hz. Resample and frame the raw single-channel recording first, then normalize each frame independently:
mean = frame.mean()scale = np.sqrt(np.mean((frame - mean) ** 2) + 1e-3)normalized = (frame - mean) / scaleRetain mean and scale to return a decoded frame to its original units: raw = decoded * scale + mean. Do not feed raw ADC or continuous values directly to encoder_int8.tflite; its calibration domain is this training-time normalized representation.
Every INT8 RVQ export reservoir-samples up to 65,536 real preprocessed validation frames, using 4,096 frames for LiteRT calibration and a disjoint 2,048-frame holdout for quantization_report.json. The release gate requires at most 1% input saturation on that holdout. A P90 reconstruction PRD target of 10% against the float32 encoder remains a quality recommendation. Worst-frame PRD remains a 15% tail-risk warning in the report, rather than a release blocker, because isolated RVQ decision-boundary crossings can be disproportionate. compressionkit golden validate-all --strict-release and the Hugging Face publisher reject INT8 RVQ releases without a passed report.
Real Browser-Demo Recordings
Section titled “Real Browser-Demo Recordings”RVQ releases carry ten 30-second real recordings in demo_recordings.npz: BIDMC PPG at 64 Hz and MIT-BIH ECG at 256 Hz. The companion manifest retains ODC-By attribution and the quality gate results. Signals are continuous source waveforms after resampling; apply the model’s normal framing and normalization before inference.
Refresh local bundles before a manual Hugging Face update:
scripts/devcontainer.sh exec -- uv run python scripts/attach_rvq_demo_recordings.py --modality allTo update only validated release candidates, repeat --experiment-id, for example --experiment-id ecg-rvq-2x --experiment-id ecg-rvq-4x.
Use --duration-seconds 120 for two-minute clips. The command updates only existing results/*/deploy/ packages, their manifests, and checksums; it does not train or re-export models.
Local Runtime Quickstart
Section titled “Local Runtime Quickstart”Use the source-checkout installation from Getting started. LiteRT executes the exported models, but importing the package also loads modules with Keras and other development dependencies; this is not a separately packaged minimal runtime.
import numpy as np
from compressionkit.runtime import RVQCodec
def synthetic_ppg_frame(frame_size: int = 320, sample_rate: int = 64) -> np.ndarray: t = np.arange(frame_size, dtype=np.float32) / np.float32(sample_rate) waveform = ( 0.55 * np.sin(2.0 * np.pi * 1.2 * t) + 0.18 * np.sin(2.0 * np.pi * 2.4 * t + 0.3) + 0.03 * np.sin(2.0 * np.pi * 0.15 * t) ) return waveform.reshape(1, 1, frame_size, 1).astype(np.float32)
codec = RVQCodec("results/ppg_rvq_64hz_04x_golden/deploy")signal = synthetic_ppg_frame()
indices = codec.encode(signal)reconstruction = codec.decode(indices)
print(indices.shape)print(reconstruction.shape)print(codec.num_levels, codec.num_embeddings, codec.embedding_dim)Use the lower-level methods when you want to separate the pipeline into explicit steps:
latent = codec.encode_latent(signal)indices = codec.quantize_latent(latent)quantized_latent = codec.dequantize_indices(indices)
if codec.has_decoder: reconstruction = codec.decode_latent(quantized_latent)This is useful when the encoder runs on-device and the decoder runs elsewhere.
HuggingFace Quickstart
Section titled “HuggingFace Quickstart”Published model repos follow the convention Ambiq/compressionkit-{modality}-{cr}x-{version}, for example Ambiq/compressionkit-ppg-4x-v1.1 or Ambiq/compressionkit-ecg-8x-v1.1.
Install the optional Hub dependency:
uv sync --extra hfThen load a codec directly from the Hub:
import numpy as np
from compressionkit.runtime import RVQCodec
def synthetic_ecg_frame(frame_size: int = 512, sample_rate: int = 256) -> np.ndarray: t = np.arange(frame_size, dtype=np.float32) / np.float32(sample_rate) waveform = ( 0.75 * np.sin(2.0 * np.pi * 1.1 * t) + 0.12 * np.sin(2.0 * np.pi * 9.0 * t) + 0.02 * np.sin(2.0 * np.pi * 0.3 * t) ) return waveform.reshape(1, 1, frame_size, 1).astype(np.float32)
codec = RVQCodec.from_pretrained("Ambiq/compressionkit-ecg-4x-v1.1")signal = synthetic_ecg_frame()
indices = codec.encode(signal)reconstruction = codec.decode(indices)The Hub staging step renames a few artifacts for distribution, but RVQCodec.from_pretrained() handles those differences for you. In particular, it maps:
config.jsontodeploy_manifest.jsonencoder_int8.tflitetoencoder.tflitedecoder_int8.tflitetodecoder.tflitesample_stimulus.npztosample_data.npz
Two-Stage Compression
Section titled “Two-Stage Compression”The two-stage path is for advanced users who want additional bitrate reduction beyond uniform RVQ token coding.
Stage 1 uses the standard RVQ codec to produce token indices. Stage 2 runs a causal entropy prior over those tokens and arithmetic-codes them into a compressed bitstream.
import numpy as np
from compressionkit.runtime import RVQCodecfrom compressionkit.runtime.prior import EntropyPriorfrom compressionkit.runtime.two_stage import TwoStageCodec
def synthetic_ppg_frame(frame_size: int = 320, sample_rate: int = 64) -> np.ndarray: t = np.arange(frame_size, dtype=np.float32) / np.float32(sample_rate) waveform = 0.6 * np.sin(2.0 * np.pi * 1.25 * t) + 0.1 * np.sin(2.0 * np.pi * 2.5 * t) return waveform.reshape(1, 1, frame_size, 1).astype(np.float32)
run_dir = "results/ppg_rvq_64hz_04x_golden"codec = RVQCodec(f"{run_dir}/deploy")prior = EntropyPrior(f"{run_dir}/deploy/prior_int8.tflite")two_stage = TwoStageCodec(codec, prior)
signal = synthetic_ppg_frame()result = two_stage.compress(signal)reconstruction = two_stage.decompress(result)
print(result.bits_per_token_uniform)print(result.bits_per_token_actual)print(result.cr_uplift)If you already have RVQ indices, you can skip the encoder and operate directly on tokens:
indices = codec.encode(signal)estimate = two_stage.estimate_bitrate(indices)compressed = two_stage.compress_indices(indices)restored_indices = two_stage.decompress_indices(compressed)Use the two-stage path when transport or storage cost is the limiting factor and you can afford the extra prior model. Use the single-stage codec when simplicity, fixed compute, or embedded deployment dominates.
SPIHT and hybrid packages use family-specific runtime loaders. Use load_codec to dispatch from package metadata, and consult the experiment page for the matching bundle target.
Platform Considerations
Section titled “Platform Considerations”
MCU / edge-device path
Section titled “MCU / edge-device path”- Keep
encoder.tfliteandcodebook.hon-device. - Add
decoder.tfliteanddecoder.honly if the product must reconstruct on the same device. - Prefer fixed frame shapes that match the exported encoder input exactly.
- Validate end-to-end quantization behavior with
sample_data.npzbefore integrating with firmware.
x86 or ARM Linux host path
Section titled “x86 or ARM Linux host path”- LiteRT can execute the exported encoder and decoder; install the package dependencies required by the Python imports.
decoder_float32.tflitesupports host reconstruction through a LiteRT interpreter rather than loading the Keras decoder model. This does not remove Keras from the package import dependencies.- HuggingFace loading is most convenient in notebooks, evaluation services, and server-side decode pipelines.
Split deployment path
Section titled “Split deployment path”- A common production split is encoder plus codebook on the wearable and decoder off-device.
- In that arrangement, transmit RVQ indices or the two-stage bitstream instead of raw windows.
- The runtime helpers (
encode_latent,quantize_latent,dequantize_indices,decode_latent) map cleanly onto that boundary.
Recommended Validation
Section titled “Recommended Validation”Before shipping a deploy package, verify these basics:
- Load the
deploy/directory withRVQCodecand confirmcodec.manifestmatches the expected frame size and latent shape. - Run
encode()anddecode()on a synthetic frame to confirm artifact completeness. - If you publish to HuggingFace, verify
RVQCodec.from_pretrained()on a clean environment. - If you ship a two-stage path, validate both
compress()anddecompress()with the exact prior artifact you plan to release.
Related Pages
Section titled “Related Pages”- See PPG models and ECG models for model-specific deployment notes.
- See Export API for the low-level export API reference.
- See Signal compression demo for a product-facing example of the codec in a demo workflow.