# Validation scorecard

A scorecard records how a reconstructed signal differs from its reference. Read it together with the dataset, sample count, preprocessing, and codec configuration. Missing measurements do not count as passing results.

## Read the metrics

| Metric | What it tells you | How to interpret it |
| --- | --- | --- |
| Compression ratio (CR) | Raw payload divided by encoded payload | State the original sample precision and whether framing overhead is included. |
| Effective CR | Payload reduction after entropy coding | Use measured bitstream size; a prior is not a fixed improvement across signals. |
| PRD | Percentage root-mean-square difference from a reference | Lower is better for a fixed reference and evaluation. |
| Faithful PRD | Error against the recorded input | Includes differences in the input's noise and artifacts. |
| Truth PRD | Error against the clean reference used by a fixture | Depends on how the reference was obtained; inspect the fixture. |
| PRDN-noise | Noise-normalized distortion | Read with the reference definition and other error metrics. A low value alone does not prove useful denoising. |
| MSE | Mean squared sample error | Sensitive to signal scale and normalization. |
| Cosine similarity | Similarity of waveform direction | Does not establish amplitude or downstream task agreement. |
| HR MAE | Mean absolute heart-rate difference, in bpm | Check the algorithm, window duration, and which frames were scored. |
| SDNN / RMSSD error | Difference in interval-variability measurements | Sensitive to detected peaks and recording duration. |
| Band error / coherence | Frequency-domain agreement | Compare using the same frequency bands and estimator. |
| Seam ratio | Boundary energy relative to window-center energy | Inspect long-recording reconstructions as well as the aggregate. |

## Agreement is different from accuracy

Applying the same algorithm to original and reconstructed signals measures agreement. It does not establish that either result is accurate against independently labeled ground truth.

A codec may preserve noise faithfully or suppress some of it. Compare error against the recorded input and, when available, a separate clean reference. Report how that reference was constructed.

## Compare like with like

Record the following alongside each comparison:

- Dataset and split, number of frames, and selection rules.
- Sample rate, frame duration, channel selection, and normalization.
- Codec version, compression ratio, and any entropy model.
- Reference signal and metric definitions.
- Noise or artifact conditions and excluded/invalid windows.
- Reconstruction overlap, stitching, and boundary handling.

The [customer evidence](https://ambiqai.github.io/compressionkit/customer-evidence/) and model pages preserve different evaluation views. Their numbers should not be combined as if they came from one run.

## ECG validation

Start with waveform error and R-peak or heart-rate agreement. If your application uses morphology or rhythm features, evaluate those specific outputs on representative recordings as well. Inspect chunk boundaries and difficult signal segments separately.

The [ECG model page](https://ambiqai.github.io/compressionkit/models/ecg/) and [ECG noise-aware tables](https://ambiqai.github.io/compressionkit/methods/cr_vs_fidelity_ecg/) show the recorded measurements. They do not establish coverage for every downstream task.

## PPG validation

Inspect pulse timing, heart-rate agreement, waveform error, and behavior under motion or baseline drift. For interval variability, use sufficiently long recordings and report the reconstruction and peak-detection procedure.

The [PPG model page](https://ambiqai.github.io/compressionkit/models/ppg/) and [PPG noise-aware tables](https://ambiqai.github.io/compressionkit/methods/cr_vs_fidelity_ppg/) show the recorded measurements. Single-channel reconstruction results do not establish multi-channel optical measurement performance.

## Signal condition buckets

Separate clean and difficult recordings rather than relying only on an overall average. Useful groups include low signal amplitude, motion, baseline drift, and boundary-adjacent samples. Use labels and thresholds that match the dataset and publish the number of samples in each group.

## Package and runtime validation

A quality score does not verify that a package loads correctly or that another runtime reproduces its output. Check manifests, file integrity, reference vectors, and encoder/decoder compatibility using the [deployment workflow](https://ambiqai.github.io/compressionkit/deployment/#recommended-validation).

Measure memory and latency in the intended environment. Define acceptance thresholds for the application before evaluating a candidate.
