Skip to content
heliaEDGE
Examples
HELIA

MiniResNet-v1

MiniResNetV1Params describes a reusable spectrogram CNN. The default topology matches ST’s one-stack MiniResNet-v1 ESC-10 checkpoint: a padded 7×7 stem, max-pooling, two residual blocks (1×1 then 3×3 convolutions), flatten and a classification head. The first block downsamples both paths with a projection; convolutions retain biases and batch normalization uses epsilon 1.001e-5.

from helia_edge.models import MiniResNetV1Params, ModelSpec, build
params = MiniResNetV1Params.model_validate({"stacks": 1, "base_filters": 64, "num_classes": 10})
model = build(ModelSpec(params=params, input_shape=(64, 50, 1)))
# The model is initialized, not pretrained. Hydrate explicitly from a local,
# verified checkpoint matching the architecture, input shape and class count:
# model.load_weights(checkpoint_path)
# model.save("classifier.keras")

Public config import/validation does not import Keras or an execution backend. A ModelSpec (parameters plus input shape) serializes with Pydantic’s JSON methods. Unknown fields and invalid dimensions are rejected, and the config is immutable. num_classes is required. One to three stacks, base channel count, pooling, head dropout and output activation are configurable; changed values are new architectures, not claims of pretrained variants. Inputs are NHWC regardless of the global image layout. Flatten needs fixed spatial dimensions; global average/max pooling supports dynamic spatial sizes. The output has num_classes scores. All layers are ordinary Keras layers, so model serialization and weight loading use standard Keras APIs. Layers default to trainable; callers control freezing when fine-tuning.

The constructor does not reset a session/seed, fetch weights, process audio or export a model. Callers own input signatures, class labels, artifact integrity, preprocessing and execution. Benchmark recipes reference their catalog’s model identity; EDGE does not maintain another asset catalog or YAML runner.

Architecture source: ST services MiniResNet-v1 at 0f6210ed. The TensorFlow/ST source notices and Apache-2.0 license are retained in helia_edge/models/licenses/miniresnet-apache-2.0.txt. The matching ST model and configuration are separately Apache-2.0 licensed by their model directory. Weights are not bundled with EDGE. The (64,50,1), ten-class default reconstruction has 126,922 parameters and 27 layers. Its trained host output parity and serialization are separate from the upstream INT8 export, model accuracy and device qualification.

The matching frontend uses 16-kHz audio, 64 mel bands, 1024-sample FFT/Hann window, hop320, 20–7500Hz, Slaney normalization, power2 and dB scaling, with centered constant padding and 50-frame patches. Exact time cleanup, dB reference, patch overlap, class ordering and clip aggregation belong to the reference recipe, not the CNN constructor. Use the pinned training configuration and loader together; raw waveform or arbitrary spectrogram normalization is not an equivalent model input. No real-audio accuracy or quantized-output equivalence is implied by host reconstruction.