# Quickstart: round-trip a golden codec

<!-- notebook-generated-page -->

Download notebookView sourceOpen in Colab

Open in Colab opens the source notebook only. Before running cells, use a Python 3.12 runtime and install compressionKIT into that runtime. See [notebook environment setup](https://ambiqai.github.io/compressionkit/examples/#notebook-environment-setup). Hosted execution has not been validated by this documentation build.

This example displays saved outputs. Building the documentation does not run training. Check dataset paths for your notebook working directory before running.

Load a published **compressionKIT** golden codec, compress and decompress a
physiological frame, and measure the true compression ratio and fidelity.

**No dataset required.** Every deploy package ships a small set of
representative, license-safe reference frames (`reference_vectors.npz`) that we
use here, so the notebook runs anywhere.

> Loading from HuggingFace needs internet access. To run fully offline, point
> `CODEC_SOURCE` at a local deploy package you built with
> `compressionkit golden run ...`.

```python title="Python"
from pathlib import Path

import matplotlib.pyplot as plt
import numpy as np
from scipy.signal import butter, sosfiltfilt

from compressionkit.evaluation.metrics import compute_signal_metrics
from compressionkit.runtime import load_codec, resolve_deploy_dir

# A published golden codec (downloads from HuggingFace): "Ambiq/compressionkit-
# {modality}-{method}-{cr}x-{version}". "-v1.0" is the current (and only)
# release track; the unsuffixed repo names used earlier in development have
# been retired.
#
#   modality : "ppg" | "ecg"
#   method   : "" (RVQ, learned/AI — omit the infix) | "spiht" (DSP-only) |
#              "hybrid" (learned denoiser + SPIHT)
#   cr       : PPG -> 2, 4, 8, 16, 32   |   ECG -> 2, 4, 8, 16, 32, 64
#              (same ladder for all three methods — all 33 combinations are
#              published under "-v1.0")
#
# All three families implement the same Codec interface, so nothing below
# this cell needs to change no matter which one you pick.
CODEC_SOURCE = "Ambiq/compressionkit-ppg-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ecg-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ppg-spiht-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ecg-spiht-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ppg-hybrid-8x-v1.0"
# CODEC_SOURCE = "Ambiq/compressionkit-ecg-hybrid-8x-v1.0"
# ... or a local deploy package you built yourself:
# CODEC_SOURCE = "results/ppg_rvq_64hz_04x_golden/deploy"
```

```text title="Saved output"
2026-07-02 23:41:19.520357: I tensorflow/core/util/port.cc:153] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
2026-07-02 23:41:19.545866: I tensorflow/core/platform/cpu_feature_guard.cc:210] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations.
To enable the following instructions: AVX2 AVX_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags.
2026-07-02 23:41:20.043290: I tensorflow/core/util/port.cc:153] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
```

```python title="Python"
codec = load_codec(CODEC_SOURCE)

print(f"name        : {codec.name}")
print(f"modality    : {codec.modality}")
print(f"sample_rate : {codec.sample_rate} Hz")
print(f"frame_size  : {codec.frame_size} samples ({codec.frame_size / codec.sample_rate:.2f} s)")
print(f"target CR   : {codec.target_cr:g}x")
```

```text title="Saved output"
Fetching 18 files: 100%|██████████| 18/18 [00:00<00:00, 6947.41it/s]
```

```text title="Saved output"
name        : ppg_rvq_64hz_08x_golden
modality    : ppg
sample_rate : 64 Hz
frame_size  : 320 samples (5.00 s)
target CR   : 8x
```

```text title="Saved output"

INFO: Created TensorFlow Lite XNNPACK delegate for CPU.
```

## Representative frames

Release packages include license-safe frames for smoke tests and demos. When
`reference_vectors.npz` is present, the notebook uses its exact runtime
conformance inputs. Otherwise it falls back to `sample_stimulus.npz` or
`sample_data.npz` from the deploy package.

```python title="Python"
deploy = Path(resolve_deploy_dir(CODEC_SOURCE))

sample_candidates = [
    ("reference_vectors.npz", ("input_frames", "inputs", "stimulus")),
    ("sample_stimulus.npz", ("input_frames", "inputs", "stimulus")),
    ("sample_data.npz", ("input_frames", "inputs", "stimulus")),
]

for filename, keys in sample_candidates:
    sample_path = deploy / filename
    if not sample_path.exists():
        continue
    sample_npz = np.load(sample_path)
    for key in keys:
        if key in sample_npz:
            sample_source = f"{filename}:{key}"
            raw_frames = sample_npz[key]
            break
    else:
        continue
    break
else:
    raise FileNotFoundError(
        f"No sample frame artifact found in {deploy}. Expected one of: "
        + ", ".join(name for name, _ in sample_candidates)
    )

raw_frames = np.asarray(raw_frames, dtype="float32")
if raw_frames.ndim == 1:
    frames = raw_frames.reshape(1, -1)
elif raw_frames.ndim == 2:
    frames = raw_frames
else:
    frames = raw_frames.reshape(raw_frames.shape[0], -1)

if frames.shape[1] != codec.frame_size:
    raise ValueError(
        f"Sample frames from {sample_source} have {frames.shape[1]} samples, "
        f"but codec.frame_size is {codec.frame_size}."
    )

print(f"Loaded {len(frames)} sample frames from {sample_source}")
```

```text title="Saved output"
Fetching 18 files: 100%|██████████| 18/18 [00:00<00:00, 6218.90it/s]
```

```text title="Saved output"
Loaded 6336 sample frames from sample_stimulus.npz:inputs
```

## Round-trip a single frame

Compression ratio is reported against a **32-bit float** sample baseline — the
toolkit's reference raw representation, which matches the headline `target CR`.

```python title="Python"
frame = frames[0]
encoded = codec.compress(frame)
recon = codec.decompress(encoded)

raw_bits = frame.size * 32  # 32-bit float baseline
true_cr = raw_bits / encoded.nbits
m = compute_signal_metrics(frame, recon)

print(f"encoded bits : {encoded.nbits}  ->  measured CR {true_cr:.2f}x (vs 32-bit float)")
print(f"PRD          : {m['prd_percent']:.2f}%")
print(f"cosine sim   : {m['cosine_similarity']:.4f}")
```

```text title="Saved output"
encoded bits : 1280  ->  measured CR 8.00x (vs 32-bit float)
PRD          : 4.75%
cosine sim   : 0.9989
```

```python title="Python"
t = np.arange(frame.size) / codec.sample_rate
plt.figure(figsize=(10, 3))
plt.plot(t, frame, label="original", lw=1.5)
plt.plot(t, recon, label="reconstruction", lw=1.2, alpha=0.85)
plt.xlabel("time (s)")
plt.ylabel("amplitude")
plt.title(f"{codec.name}  ·  PRD {m['prd_percent']:.2f}%  ·  CR {true_cr:.2f}x")
plt.legend()
plt.tight_layout()
plt.show()
```

Saved output · cell 8

## Compare against a naive baseline

How much of that fidelity is the codec actually earning? Compare against the
simplest possible scheme: an anti-alias low-pass filter, keep every `cr`-th
sample (matched to the codec's measured CR), then linearly interpolate back.
The decimated samples are stored as raw 32-bit floats — **zero entropy
coding** — so this is a floor, not a competitive baseline. It isolates how
much of a real codec's gain comes from genuine perceptual/predictive modeling
versus simply throwing samples away.

```python title="Python"
def naive_downsample_reconstruct(x: np.ndarray, sample_rate: float, cr: float) -> tuple[np.ndarray, float]:
    """Anti-alias low-pass + decimate by ~`cr`, then linearly interpolate back.

    Returns the reconstruction (same length as `x`) and the measured CR,
    assuming the decimated samples are stored as raw 32-bit floats (no
    entropy coding at all). This is a floor, not a competitive baseline.
    """
    stride = max(1, round(cr))
    nyq = 0.5 * sample_rate
    cutoff = min(0.9 * nyq / stride, 0.99 * nyq)
    sos = butter(4, cutoff / nyq, btype="low", output="sos")
    filtered = sosfiltfilt(sos, x)
    idx = np.arange(0, x.size, stride)
    decimated = filtered[idx]
    recon = np.interp(np.arange(x.size), idx, decimated)
    measured_cr = x.size / decimated.size
    return recon.astype(np.float32), measured_cr

naive_recon, naive_cr = naive_downsample_reconstruct(frame, codec.sample_rate, true_cr)
naive_metrics = compute_signal_metrics(frame, naive_recon)

print(f"naive baseline CR  : {naive_cr:.2f}x  (matched to {codec.name}'s measured CR {true_cr:.2f}x)")
print(f"naive baseline PRD : {naive_metrics['prd_percent']:.2f}%  (vs {codec.name} PRD {m['prd_percent']:.2f}%)")

plt.figure(figsize=(10, 3))
plt.plot(t, frame, label="original", lw=1.5)
plt.plot(t, recon, label=f"{codec.name} (PRD {m['prd_percent']:.2f}%)", lw=1.2, alpha=0.85)
plt.plot(
    t,
    naive_recon,
    label=f"naive filter+decimate (PRD {naive_metrics['prd_percent']:.2f}%)",
    lw=1.2,
    alpha=0.85,
    ls="--",
)
plt.xlabel("time (s)")
plt.ylabel("amplitude")
plt.title(f"Naive baseline vs {codec.name}  ·  matched CR {naive_cr:.2f}x")
plt.legend()
plt.tight_layout()
plt.show()
```

```text title="Saved output"
naive baseline CR  : 8.00x  (matched to ppg_rvq_64hz_08x_golden's measured CR 8.00x)
naive baseline PRD : 30.57%  (vs ppg_rvq_64hz_08x_golden PRD 4.75%)
```

Saved output · cell 10

## Aggregate over many frames

A single number hides the spread. Round-trip a batch and look at the
distribution of fidelity alongside the realized compression ratio.

```python title="Python"
N = min(500, len(frames))
batch = frames[:N]

total_bits = 0
prd, cos = [], []
for f in batch:
    e = codec.compress(f)
    r = codec.decompress(e)
    total_bits += e.nbits
    mm = compute_signal_metrics(f, r)
    prd.append(mm["prd_percent"])
    cos.append(mm["cosine_similarity"])

prd = np.asarray(prd)
cos = np.asarray(cos)
print(f"frames evaluated : {N}")
print(f"measured CR      : {(batch.size * 32) / total_bits:.2f}x")
print(f"PRD %            : median {np.median(prd):.2f}   mean {prd.mean():.2f}")
print(f"cosine           : median {np.median(cos):.4f}   mean {cos.mean():.4f}")
```

```text title="Saved output"
frames evaluated : 500
measured CR      : 8.00x
PRD %            : median 2.22   mean 2.64
cosine           : median 0.9998   mean 0.9994
```

```python title="Python"
plt.figure(figsize=(10, 3))
plt.hist(prd, bins=30, color="#4C78A8")
plt.axvline(np.median(prd), color="k", ls="--", label=f"median {np.median(prd):.2f}%")
plt.xlabel("PRD %")
plt.ylabel("frames")
plt.title("Per-frame fidelity distribution")
plt.legend()
plt.tight_layout()
plt.show()
```

Saved output · cell 13

## Next steps

- Swap `CODEC_SOURCE` for any tier in the [Model Zoo](https://ambiqai.github.io/compressionkit/models/) (e.g. `Ambiq/compressionkit-ecg-8x-v1.0`).
- Evaluate on **your own recordings** — see `02_evaluate_on_your_data.ipynb`.
- Reproduce a golden end-to-end: `compressionkit golden run ppg-rvq-4x`.
