Portable 1D preprocessing
Normalization1D, FirFilter and RandomGaussianNoise1D use public Keras
operations with either the TensorFlow or PyTorch backend. Importing these layers
with KERAS_BACKEND=torch does not require TensorFlow. Select the backend before
importing Keras; layers do not reset global state or switch backends.
import kerasfrom helia_edge.layers.preprocessing import Sample, Normalization1D
signal = keras.ops.ones((2, 128, 1), dtype="float32")target = keras.ops.zeros((2, 128, 1), dtype="int32")valid = keras.ops.ones((2, 128, 1), dtype="bool")sample = Sample(signals={"ecg": signal}, targets={"segmentation": target}, masks={"valid": valid})metadata = {"record_id": "example"} # retained separately by the calleroutput = Normalization1D(mean=0., variance=4.)(sample.tensor_tree())Shapes and aligned data
Section titled “Shapes and aligned data”Sample[T] and TensorPayload[T] describe two-level dictionaries of tensor
leaves. tensor_tree() creates new containers referencing the original tensors;
it does not deep-copy data, transfer devices or validate with Pydantic. The frozen
dataclass does not make tensor storage immutable. Pass its converted tree to
Keras, not the dataclass itself. Keep metadata outside the tensor tree.
Each signal is (T, C) or (B, T, C), or its channels_first counterpart.
Layers transform every entry in signals, cast only signals to their compute
dtype, and preserve target/mask values and dtypes. These three transforms do not
change temporal alignment: normalization and FIR affect signals; additive noise
affects signals only, leaving clean denoising targets unchanged. No geometric
augmentation or automatic target interpolation is implied. Noise samples are
independent across separate named signals. Caller dictionaries are not mutated.
Compatibility with tensor and dictionary calls
Section titled “Compatibility with tensor and dictionary calls”Legacy tensor calls and {"data": x, "labels": y, ...} dictionaries still work;
targets, masks and tensor-only extra fields pass through. Do not mix data
and signals. Legacy constructor options seed, auto_vectorize,
data_format, device, and Keras layer options remain accepted. These batched
transforms do not need per-example vectorization; auto_vectorize is retained
for compatibility. The default device scope remains CPU. All preprocessing now uses the unified BaseAugmentation hierarchy; see the
next-major hook and behavior migration. This is a
breaking migration for custom hooks, RNG behavior and several corrected layouts.
Filtering and random noise
Section titled “Filtering and random noise”FIR coefficients accept arrays or lists and are included in layer config.
Filtering retains the existing same-padded cross-correlation convention (it is
not a promise of SciPy lfilter or filtfilt equivalence). The same taps operate
independently on each statically known channel. forward_backward=True applies
that operation, reverses time, applies it again and reverses back. Denominator
coefficients a remain unsupported for execution and raise NotImplementedError.
Noise uses training=None by default; explicitly pass training=True to augment.
training=False or None returns signals without drawing randomness. Inference no longer advances
the noise RNG stream. Existing frozen experiments relying on the previous
inference side effect need their old version or a declared new RNG policy.
Equal seeds reproduce a fresh layer’s sequence on the same backend; cross-backend
random bit equality and restoring the exact RNG position after save/load are not
promised. Use helia_edge.models.load_model for fresh-process loading, or import
the custom classes before calling the Keras loader directly.
Pipeline integration
Section titled “Pipeline integration”With the TensorFlow backend, map the same layer in tf.data; no separate adapter
or blanket leaf casting is required:
import tensorflow as tflayer = Normalization1D(mean=0., variance=4.)dataset = tf.data.Dataset.from_tensors(sample.tensor_tree()).map(layer)Using a Torch-backed object inside a TensorFlow graph is not a supported dispatch
mode of these portable layers. Use separate backend processes if both are needed.
There is no global set_backend call or private keras.src dispatch in this core.
Grain is optional: install helia-edge[grain]. Prefer Grain batches passed to these
same layers in the consumer. For per-record worker execution, the executable
examples/preprocessing/grain_pipeline.py passes a transform to helia_edge.data.to_grain.
The transform draws a layer seed from its per-record generator, calls these layers on
explicit CPU tensors, and converts outputs to arrays. Its owned_records helper copies IPC arrays before closing the iterator;
that copy is explicit and distinct from the reference-preserving sample facade.
The example constructs small layers per record for clarity, not maximum throughput.
It creates no persistent worker RNG whose sequence depends on worker count.
Tested boundary
Section titled “Tested boundary”Local CPU checks use Keras 3.15.1, NumPy 2.3.3, TF 2.21.0/Python 3.12.5 and Torch 2.14.0/Python 3.14.7. The optional worker example uses Grain 0.2.18 without TensorFlow, with repeatable results for zero and two workers and changed results for a changed seed. CI selects the portable tests in both backend jobs; Grain checks run only when Grain is installed.
The tested two-level dictionaries work eagerly, as Functional inputs/outputs,
and through .keras save/load. This does not certify arbitrary nesting, None
leaves, ragged tensors, JAX, accelerators or export formats. Normalization/FIR
pass tf.function and Torch Dynamo capture with backend="eager", including
fullgraph. Stateful noise requires graph breaks on Torch 2.14/Keras 3.15.1;
fullgraph=True fails at Keras seed generation. No Inductor/XLA or performance
claim is made, and upstream fullgraph RNG/roll limitations are not repaired here.