to_native_fp16
PythonReturn a graph whose inputs, weights, activations and outputs are FLOAT16.
to_native_fp16(model_content: bytes) -> bytesReturn a graph whose inputs, weights, activations and outputs are FLOAT16.
The TFLite float16 optimization stores weights as FLOAT16 behind
DEQUANTIZE operators and computes in FLOAT32. This drops each
FLOAT16 -> FLOAT32 DEQUANTIZE, rewires its consumers to the FLOAT16
source, and converts every remaining FLOAT32 tensor and constant buffer to
FLOAT16. Constants outside the float16 range saturate to +/-65504, as in
TFLite’s float16 optimization. Non-float tensors are unchanged. The dropped
DEQUANTIZE outputs and operator codes no operator uses are removed, and
operator, subgraph and signature tensor indices are renumbered to match.
Buffers of removed tensors stay in the model; they are empty.
Parameters
| Name | Type | Default | Description |
|---|---|---|---|
model_content | bytes | Required | Weight-only float16 TFLite flatbuffer. |
Returns
| Value | Type | Description |
|---|---|---|
bytes | bytes | Native float16 TFLite flatbuffer. |