# Float and FP16

Compile a float32 graph and a float16 graph for the same target, and see what
the FP16 platform gate does when the hardware cannot serve it. Float support is
experimental; qualify your model, kernel build and target before deployment.

| Field | Value |
| --- | --- |
| Model | one SQRT node, built in float32 and in float16 |
| Target | `apollo510_evb` |
| Shows | Float kernel selection, the FP16 platform gate, and the ns-cmsis-nn build switches a float module needs |
| CI | Conversion and host compilation are declared in the examples CI job, both precisions |

The models are built by [make_model.py](https://github.com/AmbiqAI/helia-aot/blob/cf2246a7ad439fa38d57462118e7840e5497a685/examples/float-fp16/make_model.py). Each is a small SQRT
fixture with matching input/output dtype, so it isolates conversion's float
path without depending on a downloaded model. It is not a general test of
mixed precision, quantized weights or every float operator.

## Prerequisites

Use a repository checkout and activate its environment after `uv sync --frozen --group ci`. The model-building script imports the repository's Python helpers; an isolated CLI installation alone is not enough. See [Running examples](https://ambiqai.github.io/helia-aot/examples/).

## Run it

```sh
./run.sh
```

The two configurations, [config-float32.yaml](https://github.com/AmbiqAI/helia-aot/blob/cf2246a7ad439fa38d57462118e7840e5497a685/examples/float-fp16/config-float32.yaml) and
[config-float16.yaml](https://github.com/AmbiqAI/helia-aot/blob/cf2246a7ad439fa38d57462118e7840e5497a685/examples/float-fp16/config-float16.yaml), select the corresponding model and retain separately named output modules.
Nothing in the configuration selects a float kernel: the tensor dtypes do.

## The platform gate

`apollo510_evb` declares both a single-precision FPU and the `FP16`
capability, so both conversions resolve. Ask for float16 on a target without
it and the conversion stops before emitting anything:

```sh
helia-aot convert --model.path models/sqrt_float16.tflite --model.name sqrt_float16 \
  --module.path ./out --module.name gate --module.type cmake --platform.name apollo4p_evb
```

```text
Error: f16 kernels require a Cortex-M55-class/FP16 platform; got 'apollo4p_evb'
(cpu=cortex-m4)
  hint: Target a Cortex-M55-class/FP16 platform, or re-export the model without
float16 tensors.
```

This is a separate negative check; `run.sh` performs the two positive conversions.
FP32 does not use the FP16 platform gate, but still needs supported operator
contracts and a compatible library/toolchain build. Neither conversion nor
host compilation proves numerical SQRT behavior or device FP16 execution.

## The build switch

The float kernel objects exist only if ns-cmsis-nn was built with them, so the
switch belongs on the library rather than on the generated module. For the
CMake and NSX packagings set `ARM_NN_ENABLE_F32` or `ARM_NN_ENABLE_F16` before
adding the ns-cmsis-nn subdirectory; for Zephyr the symbols are
`CONFIG_NS_CMSIS_NN_ENABLE_F32` and `CONFIG_NS_CMSIS_NN_ENABLE_F16` in the
application's `prj.conf`. The module asks the library what it was built with,
and a mismatch fails at configure time rather than at run time.

## Next

- [Float support](https://ambiqai.github.io/helia-aot/guide/precision/) for the
  precision table, the hardware requirements and the mixed-graph rules.
- [Operator catalog](https://ambiqai.github.io/helia-aot/reference/operators/)
  for which operators have float kernels and which kernel-library floor they
  need.
