Skip to content
compressionKIT
Getting started
HELIA

Quickstart: round-trip a golden codec

Open in Colab opens the source notebook only. Before running cells, use a Python 3.12 runtime and install compressionKIT into that runtime. See notebook environment setup. Hosted execution has not been validated by this documentation build.

This example displays saved outputs. Building the documentation does not run training. Check dataset paths for your notebook working directory before running.

Load a published compressionKIT golden codec, compress and decompress a physiological frame, and measure the true compression ratio and fidelity.

No dataset required. Every deploy package ships a small set of representative, license-safe reference frames (reference_vectors.npz) that we use here, so the notebook runs anywhere.

Loading from HuggingFace needs internet access. To run fully offline, point CODEC_SOURCE at a local deploy package you built with compressionkit golden run ....

Python
from pathlib import Path
import matplotlib.pyplot as plt
import numpy as np
from scipy.signal import butter, sosfiltfilt
from compressionkit.evaluation.metrics import compute_signal_metrics
from compressionkit.runtime import load_codec, resolve_deploy_dir
# A published golden codec (downloads from HuggingFace): "Ambiq/compressionkit-
# {modality}-{method}-{cr}x-{version}". "-v1.0" is the current (and only)
# release track; the unsuffixed repo names used earlier in development have
# been retired.
#
# modality : "ppg" | "ecg"
# method : "" (RVQ, learned/AI — omit the infix) | "spiht" (DSP-only) |
# "hybrid" (learned denoiser + SPIHT)
# cr : PPG -> 2, 4, 8, 16, 32 | ECG -> 2, 4, 8, 16, 32, 64
# (same ladder for all three methods — all 33 combinations are
# published under "-v1.0")
#
# All three families implement the same Codec interface, so nothing below
# this cell needs to change no matter which one you pick.
CODEC_SOURCE = "Ambiq/compressionkit-ppg-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ecg-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ppg-spiht-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ecg-spiht-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ppg-hybrid-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ecg-hybrid-8x-v1.0"
# ... or a local deploy package you built yourself:
# CODEC_SOURCE = "results/ppg_rvq_64hz_04x_golden/deploy"
Saved output
2026-07-02 23:41:19.520357: I tensorflow/core/util/port.cc:153] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
2026-07-02 23:41:19.545866: I tensorflow/core/platform/cpu_feature_guard.cc:210] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations.
To enable the following instructions: AVX2 AVX_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags.
2026-07-02 23:41:20.043290: I tensorflow/core/util/port.cc:153] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
Python
codec = load_codec(CODEC_SOURCE)
print(f"name : {codec.name}")
print(f"modality : {codec.modality}")
print(f"sample_rate : {codec.sample_rate} Hz")
print(f"frame_size : {codec.frame_size} samples ({codec.frame_size / codec.sample_rate:.2f} s)")
print(f"target CR : {codec.target_cr:g}x")
Saved output
Fetching 18 files: 100%|██████████| 18/18 [00:00<00:00, 6947.41it/s]
Saved output
name : ppg_rvq_64hz_08x_golden
modality : ppg
sample_rate : 64 Hz
frame_size : 320 samples (5.00 s)
target CR : 8x
Saved output
INFO: Created TensorFlow Lite XNNPACK delegate for CPU.

Release packages include license-safe frames for smoke tests and demos. When reference_vectors.npz is present, the notebook uses its exact runtime conformance inputs. Otherwise it falls back to sample_stimulus.npz or sample_data.npz from the deploy package.

Python
deploy = Path(resolve_deploy_dir(CODEC_SOURCE))
sample_candidates = [
("reference_vectors.npz", ("input_frames", "inputs", "stimulus")),
("sample_stimulus.npz", ("input_frames", "inputs", "stimulus")),
("sample_data.npz", ("input_frames", "inputs", "stimulus")),
]
for filename, keys in sample_candidates:
sample_path = deploy / filename
if not sample_path.exists():
continue
sample_npz = np.load(sample_path)
for key in keys:
if key in sample_npz:
sample_source = f"{filename}:{key}"
raw_frames = sample_npz[key]
break
else:
continue
break
else:
raise FileNotFoundError(
f"No sample frame artifact found in {deploy}. Expected one of: "
+ ", ".join(name for name, _ in sample_candidates)
)
raw_frames = np.asarray(raw_frames, dtype="float32")
if raw_frames.ndim == 1:
frames = raw_frames.reshape(1, -1)
elif raw_frames.ndim == 2:
frames = raw_frames
else:
frames = raw_frames.reshape(raw_frames.shape[0], -1)
if frames.shape[1] != codec.frame_size:
raise ValueError(
f"Sample frames from {sample_source} have {frames.shape[1]} samples, "
f"but codec.frame_size is {codec.frame_size}."
)
print(f"Loaded {len(frames)} sample frames from {sample_source}")
Saved output
Fetching 18 files: 100%|██████████| 18/18 [00:00<00:00, 6218.90it/s]
Saved output
Loaded 6336 sample frames from sample_stimulus.npz:inputs

Compression ratio is reported against a 32-bit float sample baseline — the toolkit’s reference raw representation, which matches the headline target CR.

Python
frame = frames[0]
encoded = codec.compress(frame)
recon = codec.decompress(encoded)
raw_bits = frame.size * 32 # 32-bit float baseline
true_cr = raw_bits / encoded.nbits
m = compute_signal_metrics(frame, recon)
print(f"encoded bits : {encoded.nbits} -> measured CR {true_cr:.2f}x (vs 32-bit float)")
print(f"PRD : {m['prd_percent']:.2f}%")
print(f"cosine sim : {m['cosine_similarity']:.4f}")
Saved output
encoded bits : 1280 -> measured CR 8.00x (vs 32-bit float)
PRD : 4.75%
cosine sim : 0.9989
Python
t = np.arange(frame.size) / codec.sample_rate
plt.figure(figsize=(10, 3))
plt.plot(t, frame, label="original", lw=1.5)
plt.plot(t, recon, label="reconstruction", lw=1.2, alpha=0.85)
plt.xlabel("time (s)")
plt.ylabel("amplitude")
plt.title(f"{codec.name} · PRD {m['prd_percent']:.2f}% · CR {true_cr:.2f}x")
plt.legend()
plt.tight_layout()
plt.show()
PPG waveform comparison: original and codec reconstruction amplitude over time in seconds.
Saved output · cell 8

How much of that fidelity is the codec actually earning? Compare against the simplest possible scheme: an anti-alias low-pass filter, keep every cr-th sample (matched to the codec’s measured CR), then linearly interpolate back. The decimated samples are stored as raw 32-bit floats — zero entropy coding — so this is a floor, not a competitive baseline. It isolates how much of a real codec’s gain comes from genuine perceptual/predictive modeling versus simply throwing samples away.

Python
def naive_downsample_reconstruct(x: np.ndarray, sample_rate: float, cr: float) -> tuple[np.ndarray, float]:
"""Anti-alias low-pass + decimate by ~`cr`, then linearly interpolate back.
Returns the reconstruction (same length as `x`) and the measured CR,
assuming the decimated samples are stored as raw 32-bit floats (no
entropy coding at all). This is a floor, not a competitive baseline.
"""
stride = max(1, round(cr))
nyq = 0.5 * sample_rate
cutoff = min(0.9 * nyq / stride, 0.99 * nyq)
sos = butter(4, cutoff / nyq, btype="low", output="sos")
filtered = sosfiltfilt(sos, x)
idx = np.arange(0, x.size, stride)
decimated = filtered[idx]
recon = np.interp(np.arange(x.size), idx, decimated)
measured_cr = x.size / decimated.size
return recon.astype(np.float32), measured_cr
naive_recon, naive_cr = naive_downsample_reconstruct(frame, codec.sample_rate, true_cr)
naive_metrics = compute_signal_metrics(frame, naive_recon)
print(f"naive baseline CR : {naive_cr:.2f}x (matched to {codec.name}'s measured CR {true_cr:.2f}x)")
print(f"naive baseline PRD : {naive_metrics['prd_percent']:.2f}% (vs {codec.name} PRD {m['prd_percent']:.2f}%)")
plt.figure(figsize=(10, 3))
plt.plot(t, frame, label="original", lw=1.5)
plt.plot(t, recon, label=f"{codec.name} (PRD {m['prd_percent']:.2f}%)", lw=1.2, alpha=0.85)
plt.plot(
t,
naive_recon,
label=f"naive filter+decimate (PRD {naive_metrics['prd_percent']:.2f}%)",
lw=1.2,
alpha=0.85,
ls="--",
)
plt.xlabel("time (s)")
plt.ylabel("amplitude")
plt.title(f"Naive baseline vs {codec.name} · matched CR {naive_cr:.2f}x")
plt.legend()
plt.tight_layout()
plt.show()
Saved output
naive baseline CR : 8.00x (matched to ppg_rvq_64hz_08x_golden's measured CR 8.00x)
naive baseline PRD : 30.57% (vs ppg_rvq_64hz_08x_golden PRD 4.75%)
PPG waveform comparison: original, codec reconstruction, and low-pass filter plus decimation baseline at matched measured compression ratios. Horizontal axis: time in seconds; vertical axis: amplitude.
Saved output · cell 10

A single number hides the spread. Round-trip a batch and look at the distribution of fidelity alongside the realized compression ratio.

Python
N = min(500, len(frames))
batch = frames[:N]
total_bits = 0
prd, cos = [], []
for f in batch:
e = codec.compress(f)
r = codec.decompress(e)
total_bits += e.nbits
mm = compute_signal_metrics(f, r)
prd.append(mm["prd_percent"])
cos.append(mm["cosine_similarity"])
prd = np.asarray(prd)
cos = np.asarray(cos)
print(f"frames evaluated : {N}")
print(f"measured CR : {(batch.size * 32) / total_bits:.2f}x")
print(f"PRD % : median {np.median(prd):.2f} mean {prd.mean():.2f}")
print(f"cosine : median {np.median(cos):.4f} mean {cos.mean():.4f}")
Saved output
frames evaluated : 500
measured CR : 8.00x
PRD % : median 2.22 mean 2.64
cosine : median 0.9998 mean 0.9994
Python
plt.figure(figsize=(10, 3))
plt.hist(prd, bins=30, color="#4C78A8")
plt.axvline(np.median(prd), color="k", ls="--", label=f"median {np.median(prd):.2f}%")
plt.xlabel("PRD %")
plt.ylabel("frames")
plt.title("Per-frame fidelity distribution")
plt.legend()
plt.tight_layout()
plt.show()
Histogram of per-frame PPG percentage root-mean-square difference (PRD). Horizontal axis: PRD percent; vertical axis: frame count. A dashed line marks the median.
Saved output · cell 13
  • Swap CODEC_SOURCE for any tier in the Model Zoo (e.g. Ambiq/compressionkit-ecg-8x-v1.0).
  • Evaluate on your own recordings — see 02_evaluate_on_your_data.ipynb.
  • Reproduce a golden end-to-end: compressionkit golden run ppg-rvq-4x.