Skip to content
heliaEDGE
Reference
HELIA

fp16

Rewrite a weight-only float16 TFLite model into a native float16 graph.

Machine-readable model

  • to_native_fp16functionReturn a graph whose inputs, weights, activations and outputs are FLOAT16.
function

Return a graph whose inputs, weights, activations and outputs are FLOAT16.

helia_edge/export/fp16.py:13

to_native_fp16(model_content: bytes) -> bytes

Return a graph whose inputs, weights, activations and outputs are FLOAT16.

The TFLite float16 optimization stores weights as FLOAT16 behind DEQUANTIZE operators and computes in FLOAT32. This drops each FLOAT16 -> FLOAT32 DEQUANTIZE, rewires its consumers to the FLOAT16 source, and converts every remaining FLOAT32 tensor and constant buffer to FLOAT16. Constants outside the float16 range saturate to +/-65504, as in TFLite’s float16 optimization. Non-float tensors are unchanged. The dropped DEQUANTIZE outputs and operator codes no operator uses are removed, and operator, subgraph and signature tensor indices are renumbered to match. Buffers of removed tensors stay in the model; they are empty.

Parameters of to_native_fp16
NameTypeDefaultDescription
model_contentbytesRequiredWeight-only float16 TFLite flatbuffer.
Returns of to_native_fp16
ValueTypeDescription
bytesbytesNative float16 TFLite flatbuffer.