Metrics and evaluation
Choose metrics according to the task and the meaning of the model outputs. Accuracy, signal reconstruction error and model compute describe different properties; none substitutes for the others.
Select a metric
Section titled “Select a metric”| Task | API | Interpretation |
|---|---|---|
| Classification with extra spatial or temporal axes | MultiF1Score | Collapses leading dimensions and treats the last dimension as classes. |
| Classification error analysis | ConfusionMatrix | Inspect errors between classes, alongside aggregate scores. |
| Signal reconstruction | Snr | Signal-to-noise ratio in dB, with the reference signal as y_true. |
| Reconstruction distortion | PRD and TruePRD | Check the denominator and normalization definition before comparing values. |
| Confidence filtering | Threshold helpers | Select predictions according to their probability threshold. |
| Static compute analysis | get_flops | TensorFlow-based operation counting; it does not measure device latency. |
For a reconstruction task, the standalone SNR metric can be updated from held-out reference and predicted signals:
from helia_edge.metrics import Snr
metric = Snr()metric.update_state(reference_signals, predicted_signals)print(metric.result())metric.reset_state()The arrays must represent aligned samples with the same signal interpretation. Consult each metric’s API for weighting and reduction behavior.
Plot the result
Section titled “Plot the result”In a source checkout, run uv sync --extra plotting alongside your backend extra to visualize training progress and classification results.
Keep task metrics from held-out data separate from conversion checks. After export, compare converted predictions with the source model, then evaluate the converted model on your task’s validation data.