Quickstart: round-trip a golden codec
Open in Colab opens the source notebook only. Before running cells, use a Python 3.12 runtime and install compressionKIT into that runtime. See notebook environment setup. Hosted execution has not been validated by this documentation build.
This example displays saved outputs. Building the documentation does not run training. Check dataset paths for your notebook working directory before running.
Load a published compressionKIT golden codec, compress and decompress a physiological frame, and measure the true compression ratio and fidelity.
No dataset required. Every deploy package ships a small set of
representative, license-safe reference frames (reference_vectors.npz) that we
use here, so the notebook runs anywhere.
Loading from HuggingFace needs internet access. To run fully offline, point
CODEC_SOURCEat a local deploy package you built withcompressionkit golden run ....
from pathlib import Path
import matplotlib.pyplot as pltimport numpy as npfrom scipy.signal import butter, sosfiltfilt
from compressionkit.evaluation.metrics import compute_signal_metricsfrom compressionkit.runtime import load_codec, resolve_deploy_dir
# A published golden codec (downloads from HuggingFace): "Ambiq/compressionkit-# {modality}-{method}-{cr}x-{version}". "-v1.0" is the current (and only)# release track; the unsuffixed repo names used earlier in development have# been retired.## modality : "ppg" | "ecg"# method : "" (RVQ, learned/AI — omit the infix) | "spiht" (DSP-only) |# "hybrid" (learned denoiser + SPIHT)# cr : PPG -> 2, 4, 8, 16, 32 | ECG -> 2, 4, 8, 16, 32, 64# (same ladder for all three methods — all 33 combinations are# published under "-v1.0")## All three families implement the same Codec interface, so nothing below# this cell needs to change no matter which one you pick.CODEC_SOURCE = "Ambiq/compressionkit-ppg-8x-v1.0"# CODEC_SOURCE = "Ambiq/compressionkit-ecg-8x-v1.0"# CODEC_SOURCE = "Ambiq/compressionkit-ppg-spiht-8x-v1.0"# CODEC_SOURCE = "Ambiq/compressionkit-ecg-spiht-8x-v1.0"# CODEC_SOURCE = "Ambiq/compressionkit-ppg-hybrid-8x-v1.0"# CODEC_SOURCE = "Ambiq/compressionkit-ecg-hybrid-8x-v1.0"# ... or a local deploy package you built yourself:# CODEC_SOURCE = "results/ppg_rvq_64hz_04x_golden/deploy"2026-07-02 23:41:19.520357: I tensorflow/core/util/port.cc:153] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.2026-07-02 23:41:19.545866: I tensorflow/core/platform/cpu_feature_guard.cc:210] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations.To enable the following instructions: AVX2 AVX_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags.2026-07-02 23:41:20.043290: I tensorflow/core/util/port.cc:153] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.codec = load_codec(CODEC_SOURCE)
print(f"name : {codec.name}")print(f"modality : {codec.modality}")print(f"sample_rate : {codec.sample_rate} Hz")print(f"frame_size : {codec.frame_size} samples ({codec.frame_size / codec.sample_rate:.2f} s)")print(f"target CR : {codec.target_cr:g}x")Fetching 18 files: 100%|██████████| 18/18 [00:00<00:00, 6947.41it/s]name : ppg_rvq_64hz_08x_goldenmodality : ppgsample_rate : 64 Hzframe_size : 320 samples (5.00 s)target CR : 8xINFO: Created TensorFlow Lite XNNPACK delegate for CPU.Representative frames
Section titled “Representative frames”Release packages include license-safe frames for smoke tests and demos. When
reference_vectors.npz is present, the notebook uses its exact runtime
conformance inputs. Otherwise it falls back to sample_stimulus.npz or
sample_data.npz from the deploy package.
deploy = Path(resolve_deploy_dir(CODEC_SOURCE))
sample_candidates = [ ("reference_vectors.npz", ("input_frames", "inputs", "stimulus")), ("sample_stimulus.npz", ("input_frames", "inputs", "stimulus")), ("sample_data.npz", ("input_frames", "inputs", "stimulus")),]
for filename, keys in sample_candidates: sample_path = deploy / filename if not sample_path.exists(): continue sample_npz = np.load(sample_path) for key in keys: if key in sample_npz: sample_source = f"{filename}:{key}" raw_frames = sample_npz[key] break else: continue breakelse: raise FileNotFoundError( f"No sample frame artifact found in {deploy}. Expected one of: " + ", ".join(name for name, _ in sample_candidates) )
raw_frames = np.asarray(raw_frames, dtype="float32")if raw_frames.ndim == 1: frames = raw_frames.reshape(1, -1)elif raw_frames.ndim == 2: frames = raw_frameselse: frames = raw_frames.reshape(raw_frames.shape[0], -1)
if frames.shape[1] != codec.frame_size: raise ValueError( f"Sample frames from {sample_source} have {frames.shape[1]} samples, " f"but codec.frame_size is {codec.frame_size}." )
print(f"Loaded {len(frames)} sample frames from {sample_source}")Fetching 18 files: 100%|██████████| 18/18 [00:00<00:00, 6218.90it/s]Loaded 6336 sample frames from sample_stimulus.npz:inputsRound-trip a single frame
Section titled “Round-trip a single frame”Compression ratio is reported against a 32-bit float sample baseline — the
toolkit’s reference raw representation, which matches the headline target CR.
frame = frames[0]encoded = codec.compress(frame)recon = codec.decompress(encoded)
raw_bits = frame.size * 32 # 32-bit float baselinetrue_cr = raw_bits / encoded.nbitsm = compute_signal_metrics(frame, recon)
print(f"encoded bits : {encoded.nbits} -> measured CR {true_cr:.2f}x (vs 32-bit float)")print(f"PRD : {m['prd_percent']:.2f}%")print(f"cosine sim : {m['cosine_similarity']:.4f}")encoded bits : 1280 -> measured CR 8.00x (vs 32-bit float)PRD : 4.75%cosine sim : 0.9989t = np.arange(frame.size) / codec.sample_rateplt.figure(figsize=(10, 3))plt.plot(t, frame, label="original", lw=1.5)plt.plot(t, recon, label="reconstruction", lw=1.2, alpha=0.85)plt.xlabel("time (s)")plt.ylabel("amplitude")plt.title(f"{codec.name} · PRD {m['prd_percent']:.2f}% · CR {true_cr:.2f}x")plt.legend()plt.tight_layout()plt.show()
Compare against a naive baseline
Section titled “Compare against a naive baseline”How much of that fidelity is the codec actually earning? Compare against the
simplest possible scheme: an anti-alias low-pass filter, keep every cr-th
sample (matched to the codec’s measured CR), then linearly interpolate back.
The decimated samples are stored as raw 32-bit floats — zero entropy
coding — so this is a floor, not a competitive baseline. It isolates how
much of a real codec’s gain comes from genuine perceptual/predictive modeling
versus simply throwing samples away.
def naive_downsample_reconstruct(x: np.ndarray, sample_rate: float, cr: float) -> tuple[np.ndarray, float]: """Anti-alias low-pass + decimate by ~`cr`, then linearly interpolate back.
Returns the reconstruction (same length as `x`) and the measured CR, assuming the decimated samples are stored as raw 32-bit floats (no entropy coding at all). This is a floor, not a competitive baseline. """ stride = max(1, round(cr)) nyq = 0.5 * sample_rate cutoff = min(0.9 * nyq / stride, 0.99 * nyq) sos = butter(4, cutoff / nyq, btype="low", output="sos") filtered = sosfiltfilt(sos, x) idx = np.arange(0, x.size, stride) decimated = filtered[idx] recon = np.interp(np.arange(x.size), idx, decimated) measured_cr = x.size / decimated.size return recon.astype(np.float32), measured_cr
naive_recon, naive_cr = naive_downsample_reconstruct(frame, codec.sample_rate, true_cr)naive_metrics = compute_signal_metrics(frame, naive_recon)
print(f"naive baseline CR : {naive_cr:.2f}x (matched to {codec.name}'s measured CR {true_cr:.2f}x)")print(f"naive baseline PRD : {naive_metrics['prd_percent']:.2f}% (vs {codec.name} PRD {m['prd_percent']:.2f}%)")
plt.figure(figsize=(10, 3))plt.plot(t, frame, label="original", lw=1.5)plt.plot(t, recon, label=f"{codec.name} (PRD {m['prd_percent']:.2f}%)", lw=1.2, alpha=0.85)plt.plot( t, naive_recon, label=f"naive filter+decimate (PRD {naive_metrics['prd_percent']:.2f}%)", lw=1.2, alpha=0.85, ls="--",)plt.xlabel("time (s)")plt.ylabel("amplitude")plt.title(f"Naive baseline vs {codec.name} · matched CR {naive_cr:.2f}x")plt.legend()plt.tight_layout()plt.show()naive baseline CR : 8.00x (matched to ppg_rvq_64hz_08x_golden's measured CR 8.00x)naive baseline PRD : 30.57% (vs ppg_rvq_64hz_08x_golden PRD 4.75%)
Aggregate over many frames
Section titled “Aggregate over many frames”A single number hides the spread. Round-trip a batch and look at the distribution of fidelity alongside the realized compression ratio.
N = min(500, len(frames))batch = frames[:N]
total_bits = 0prd, cos = [], []for f in batch: e = codec.compress(f) r = codec.decompress(e) total_bits += e.nbits mm = compute_signal_metrics(f, r) prd.append(mm["prd_percent"]) cos.append(mm["cosine_similarity"])
prd = np.asarray(prd)cos = np.asarray(cos)print(f"frames evaluated : {N}")print(f"measured CR : {(batch.size * 32) / total_bits:.2f}x")print(f"PRD % : median {np.median(prd):.2f} mean {prd.mean():.2f}")print(f"cosine : median {np.median(cos):.4f} mean {cos.mean():.4f}")frames evaluated : 500measured CR : 8.00xPRD % : median 2.22 mean 2.64cosine : median 0.9998 mean 0.9994plt.figure(figsize=(10, 3))plt.hist(prd, bins=30, color="#4C78A8")plt.axvline(np.median(prd), color="k", ls="--", label=f"median {np.median(prd):.2f}%")plt.xlabel("PRD %")plt.ylabel("frames")plt.title("Per-frame fidelity distribution")plt.legend()plt.tight_layout()plt.show()
Next steps
Section titled “Next steps”- Swap
CODEC_SOURCEfor any tier in the Model Zoo (e.g.Ambiq/compressionkit-ecg-8x-v1.0). - Evaluate on your own recordings — see
02_evaluate_on_your_data.ipynb. - Reproduce a golden end-to-end:
compressionkit golden run ppg-rvq-4x.