Skip to content
heliaEDGE
User guide
HELIA

Portable 1D preprocessing

Normalization1D, FirFilter and RandomGaussianNoise1D use public Keras operations with either the TensorFlow or PyTorch backend. Importing these layers with KERAS_BACKEND=torch does not require TensorFlow. Select the backend before importing Keras; layers do not reset global state or switch backends.

import keras
from helia_edge.layers.preprocessing import Sample, Normalization1D
signal = keras.ops.ones((2, 128, 1), dtype="float32")
target = keras.ops.zeros((2, 128, 1), dtype="int32")
valid = keras.ops.ones((2, 128, 1), dtype="bool")
sample = Sample(signals={"ecg": signal},
targets={"segmentation": target}, masks={"valid": valid})
metadata = {"record_id": "example"} # retained separately by the caller
output = Normalization1D(mean=0., variance=4.)(sample.tensor_tree())

Sample[T] and TensorPayload[T] describe two-level dictionaries of tensor leaves. tensor_tree() creates new containers referencing the original tensors; it does not deep-copy data, transfer devices or validate with Pydantic. The frozen dataclass does not make tensor storage immutable. Pass its converted tree to Keras, not the dataclass itself. Keep metadata outside the tensor tree.

Each signal is (T, C) or (B, T, C), or its channels_first counterpart. Layers transform every entry in signals, cast only signals to their compute dtype, and preserve target/mask values and dtypes. These three transforms do not change temporal alignment: normalization and FIR affect signals; additive noise affects signals only, leaving clean denoising targets unchanged. No geometric augmentation or automatic target interpolation is implied. Noise samples are independent across separate named signals. Caller dictionaries are not mutated.

Compatibility with tensor and dictionary calls

Section titled “Compatibility with tensor and dictionary calls”

Legacy tensor calls and {"data": x, "labels": y, ...} dictionaries still work; targets, masks and tensor-only extra fields pass through. Do not mix data and signals. Legacy constructor options seed, auto_vectorize, data_format, device, and Keras layer options remain accepted. These batched transforms do not need per-example vectorization; auto_vectorize is retained for compatibility. The default device scope remains CPU. All preprocessing now uses the unified BaseAugmentation hierarchy; see the next-major hook and behavior migration. This is a breaking migration for custom hooks, RNG behavior and several corrected layouts.

FIR coefficients accept arrays or lists and are included in layer config. Filtering retains the existing same-padded cross-correlation convention (it is not a promise of SciPy lfilter or filtfilt equivalence). The same taps operate independently on each statically known channel. forward_backward=True applies that operation, reverses time, applies it again and reverses back. Denominator coefficients a remain unsupported for execution and raise NotImplementedError.

Noise uses training=None by default; explicitly pass training=True to augment. training=False or None returns signals without drawing randomness. Inference no longer advances the noise RNG stream. Existing frozen experiments relying on the previous inference side effect need their old version or a declared new RNG policy. Equal seeds reproduce a fresh layer’s sequence on the same backend; cross-backend random bit equality and restoring the exact RNG position after save/load are not promised. Use helia_edge.models.load_model for fresh-process loading, or import the custom classes before calling the Keras loader directly.

With the TensorFlow backend, map the same layer in tf.data; no separate adapter or blanket leaf casting is required:

import tensorflow as tf
layer = Normalization1D(mean=0., variance=4.)
dataset = tf.data.Dataset.from_tensors(sample.tensor_tree()).map(layer)

Using a Torch-backed object inside a TensorFlow graph is not a supported dispatch mode of these portable layers. Use separate backend processes if both are needed. There is no global set_backend call or private keras.src dispatch in this core.

Grain is optional: install helia-edge[grain]. Prefer Grain batches passed to these same layers in the consumer. For per-record worker execution, the executable examples/preprocessing/grain_pipeline.py passes a transform to helia_edge.data.to_grain. The transform draws a layer seed from its per-record generator, calls these layers on explicit CPU tensors, and converts outputs to arrays. Its owned_records helper copies IPC arrays before closing the iterator; that copy is explicit and distinct from the reference-preserving sample facade. The example constructs small layers per record for clarity, not maximum throughput. It creates no persistent worker RNG whose sequence depends on worker count.

Local CPU checks use Keras 3.15.1, NumPy 2.3.3, TF 2.21.0/Python 3.12.5 and Torch 2.14.0/Python 3.14.7. The optional worker example uses Grain 0.2.18 without TensorFlow, with repeatable results for zero and two workers and changed results for a changed seed. CI selects the portable tests in both backend jobs; Grain checks run only when Grain is installed.

The tested two-level dictionaries work eagerly, as Functional inputs/outputs, and through .keras save/load. This does not certify arbitrary nesting, None leaves, ragged tensors, JAX, accelerators or export formats. Normalization/FIR pass tf.function and Torch Dynamo capture with backend="eager", including fullgraph. Stateful noise requires graph breaks on Torch 2.14/Keras 3.15.1; fullgraph=True fails at Keras seed generation. No Inductor/XLA or performance claim is made, and upstream fullgraph RNG/roll limitations are not repaired here.