Skip to content
compressionKIT
Getting started
HELIA

Dataset Setup

compressionKIT separates using a published codec from training or evaluating on real data:

  • Running a golden codec (load from HuggingFace, compress/decompress a frame) needs no dataset at all — see Getting Started and the example notebooks.
  • Reproducing a golden run (training) or evaluating on physiological recordings needs the raw datasets below.

If you only want to try a codec on your own signals, skip straight to Bring your own data or the synthetic generators.

Run the examples from the repository root, where the golden workflows expect a datasets/ directory. Configuration fields and scripts that use the shared dataset-root default read COMPRESSIONKIT_DATASETS_DIR:

Terminal window
export COMPRESSIONKIT_DATASETS_DIR=/path/to/your/datasets
Shared default resolutionValue
1. Environment variableCOMPRESSIONKIT_DATASETS_DIR
2. Fallback defaultdatasets (relative to the working directory)

This variable does not override explicit YAML paths or the golden runner’s dataset preflight. For example, the PPG RVQ goldens explicitly set data.unified_cache.cache_root: datasets/ppg_cache_strict_sanitize. The runner’s --datasets-root flag changes the preflight root and supplies data paths to SPIHT/hybrid builds, but does not rewrite RVQ training configs. When using separate locations, keep the preflight root, training configuration, and cache-build output paths consistent.

datasets/
├── ptbxl/ # ECG · PTB-XL (converted to .h5)
│ └── *.h5
├── bidmc/ # PPG · BIDMC
│ └── *.h5
├── ppg_dalia/ # PPG · PPG-DaLiA
│ └── *.h5
├── wesad/ # PPG · WESAD
│ └── *.h5
├── butppg/ # PPG · BUT PPG
│ └── *.h5
├── mesa-commercial-use/ # PPG · MESA (NSRR, restricted)
│ └── polysomnography/edfs/*.edf
├── ppg_cache_strict_sanitize/ # Generated PPG v1 golden cache (built once)
│ ├── bidmc/{train,val}.tfrecord
│ ├── butppg/{train,val}.tfrecord
│ ├── ppg_dalia/{train,val}.tfrecord
│ └── wesad/{train,val}.tfrecord
├── ppg_cache/ # Optional ad hoc PPG caches
└── ecg_tfrecord_cache/ # Generated TFRecord cache (built once)

PTB-XL is released under CC BY 4.0 and may be redistributed. The loader ships a one-time downloader that fetches and converts records to HDF5:

from compressionkit.datasets import PtbxlDataset
ds = PtbxlDataset(path="datasets/ptbxl")
ds.download() # one-time: fetch + convert to .h5
sig = ds.load_signal(patient_id=1) # (12, 5000) float32 @ 500 Hz

Golden ECG configs read 256 Hz, lead II windows via data.dataset_glob: ptbxl/*.h5.

The PPG golden codecs train on a unified mixture of open PPG datasets. MESA is NSRR-restricted and must not be redistributed — request access through the NSRR. The remaining sources are openly available:

SlugDatasetFormatNotes
mesaMESA (PSG).edfNSRR restricted; not used in published goldens
bidmcBIDMC.h5Open
ppg_daliaPPG-DaLiA.h5Open
wesadWESAD.h5Open
butppgBUT PPG.h5Open; quality-filtered

Install the ingestion helpers, then fetch or convert each open source into the canonical datasets/<slug>/*.h5 layout. The download scripts accept --limit for a smoke test before a full pull.

Terminal window
uv sync --group ingest
# Optional smoke tests first.
uv run python scripts/datasets/download_bidmc.py --limit 3
uv run python scripts/datasets/download_butppg.py --limit 5
uv run python scripts/datasets/download_ppg_dalia.py --limit 2
uv run python scripts/datasets/download_wesad.py --limit 2
# Full source preparation for the v1 PPG golden cache.
uv run python scripts/datasets/download_bidmc.py
uv run python scripts/datasets/download_butppg.py
uv run python scripts/datasets/download_ppg_dalia.py
uv run python scripts/datasets/download_wesad.py

If a source is already present, the scripts skip work or reuse existing files where possible. Keep each source under the same root selected by COMPRESSIONKIT_DATASETS_DIR.

Each source is cached once into TFRecords, then any combination can be mixed during training:

Terminal window
# Build the exact open-source cache expected by the v1 PPG goldens.
uv run python scripts/build_ppg_cache.py \
--sources bidmc butppg ppg_dalia wesad \
--cache-root datasets/ppg_cache_strict_sanitize
# Custom locations: set both the source root and cache output explicitly.
uv run python scripts/build_ppg_cache.py --sources bidmc \
--datasets-root /data/datasets \
--cache-root /data/ppg_cache_strict_sanitize

The golden runner checks for train.tfrecord, val.tfrecord, and metadata.json under each required source directory. If the cache is missing, compressionkit golden run ppg-rvq-8x fails early with the same build command instead of failing deep in training.

Load a one-dimensional waveform, resample it to the codec sample rate, and follow its preprocessing contract. This example assumes your_signal is already a float32 array with those properties:

import numpy as np
from compressionkit.runtime import load_codec
codec = load_codec("Ambiq/compressionkit-ppg-4x-v1.1")
fs, n = codec.sample_rate, codec.frame_size # PPG: 64 Hz
# your_signal: 1-D float32 sampled at `fs`
frames = your_signal[: len(your_signal) // n * n].reshape(-1, n)
recon = np.concatenate([codec.decompress(codec.compress(f)) for f in frames])

This snippet drops an incomplete final frame and does not implement overlap or stitching. See the evaluation notebook for a fuller workflow.

For demos, CI, and quick sanity checks, compressionKIT ships analytical PPG/ECG generators (no datasets, fully license-safe):

from compressionkit.synthetic import ppg_dynamical, ecg_mcsharry, add_noise, NoiseSpec
clean = ppg_dynamical(duration_s=30.0, sample_rate=64.0, hr_mean=72.0)
ecg = ecg_mcsharry(duration_s=10.0, sample_rate=256.0, hr_mean=60.0)
noisy, noise = add_noise(clean, sample_rate=64.0, snr_db=15.0)

These pair a ground-truth-clean waveform with a controlled noise harness, which is exactly what the example notebooks use so they run anywhere.

DatasetModalityLicenseRedistribute?
PTB-XLECGCC BY 4.0Yes
MESAPPGNSRR DUANo
BIDMCPPGOpen (PhysioNet)Per source terms
PPG-DaLiAPPGOpenPer source terms
WESADPPGOpenPer source terms
BUT PPGPPGOpenPer source terms
Synthetic (physiokit)PPG/ECGGeneratedYes

Always confirm each source’s terms before redistributing derived artifacts.