FP16 and FP32 HELIA Kernels
heliaRT can dispatch selected TensorFlow Lite Micro floating-point operators to the optimized FP32 and FP16 kernels supplied by AmbiqAI/ns-cmsis-nn. The heliaRT operator wrappers and the linked ns-cmsis-nn library must agree on which float APIs are available.
Feature contract
HELIA float builds require ns-cmsis-nn v7.33.0 or later for the
SPLIT, PACK, UNPACK, and FILL adapters. These operators support
FP16 and call the corresponding CORE APIs for FP32 when enabled. FP32
still uses the existing reference implementation when ARM_NN_ENABLE_F32=0;
FP16 is rejected during AllocateTensors() when ARM_NN_ENABLE_F16=0.
Update separately supplied CORE source targets or archives along with RT.
These four operations preserve storage bits, including NaN payloads, signed zero and subnormals. They accept negative axes after normalization, scalar PACK inputs and FILL outputs, and valid empty tensors. Float shapes, counts, and types are checked during preparation. Shapes must fit the CORE int32 element-count and target address-size limits; pointer arrays and shape metadata use the tensor arena rather than fixed rank or tensor-count limits. Empty outputs perform no data access. This adoption makes no performance guarantee.
ns-cmsis-nn v7.28.0 or later exports these definitions from its CMake target; from v7.32.0 they are also the only names it accepts as CMake inputs, including on the NSX path:
They serve two purposes:
- ns-cmsis-nn uses the corresponding CMake options to select the FP32 and FP16 kernel sources compiled into its library.
- heliaRT uses the exported definitions to compile only calls supported by that library.
When both projects are built from source through CMake, the resolved ns-cmsis-nn target is authoritative. heliaRT adopts its exported values so the operator wrappers and kernel archive cannot be compiled with conflicting float definitions.
Source of truth by build system
| Build system | Float-feature source of truth | How alignment works |
|---|---|---|
| Make | tensorflow/lite/micro/tools/make/ext_libs/helia.inc |
Make compiles the heliaRT C++ wrappers and selected ns-cmsis-nn C sources directly into the same libtensorflow-microlite.a. helia.inc selects the FP sources and adds matching ARM_NN_ENABLE_F32/F16 definitions to both C and C++ compilation. |
| Zephyr | Final Kconfig values | Kconfig resolves before either module's CMake file runs. Both heliaRT and ns-cmsis-nn consume the same CONFIG_NS_CMSIS_NN_ENABLE_F32/F16 values while compiling their respective libraries. |
| NSX/CMake | The resolved ns-cmsis-nn target | ns-cmsis-nn selects its sources and exports ARM_NN_ENABLE_F32/F16. heliaRT reads and adopts those definitions, then links ns-cmsis-nn transitively into the application. |
An ordinary CMake static library does not physically absorb another static
library. In source NSX/CMake builds, helia_rt::helia carries
nsx::cmsis_nn in its public link interface, so the final application link
includes both archives. The Make release builder is different: it compiles both
projects' objects into one combined archive.
Important
Resolve the CMake feature options during configuration, before
add_subdirectory() processes ns-cmsis-nn — either as -D arguments on the
cmake command line, or as a set(... CACHE BOOL "" FORCE) earlier in the
parent CMakeLists.txt. Adding -DARM_NN_ENABLE_F16=1 to compiler flags is
not sufficient because source selection has already happened.
Hardware support
The release archives support the following matrix:
| Target | Optimized FP32 | Optimized FP16 | Requirement |
|---|---|---|---|
cortex-m4+fp |
Yes | No | Cortex-M4 single-precision FPU |
cortex-m55 |
Yes | Yes | Armv8.1-M with MVE floating point (MVEF/Helium) |
FP16 requires the compiler's float16_t support and an MVEF-capable target.
The Make build therefore enables FP16 only for TARGET_ARCH=cortex-m55.
Do not use the M55 FP16 archive on a Cortex-M55 configuration where floating
point MVE is disabled.
The published static-library matrix does not provide optimized floating-point kernels for Cortex-M0 or Cortex-M4 builds without an FPU.
Runtime behavior
The selected backend is fixed at build time, but each operator still validates whether its tensor types, shapes, layout, padding, activation, and other parameters are supported by the optimized kernel.
- FP32 operators generally fall back to the TFLM reference implementation when the optimized API is disabled or rejects the configuration.
- Many FP16 operators do not have a TFLM FP16 reference implementation. Where
the limitation is knowable at graph preparation (FP16 disabled at build time,
broadcast
ADD/MUL, non-4DPAD, groupedCONV_2D, non-unit-betaSOFTMAX), the operator failsAllocateTensors()with a logged message; configurations the optimized kernel rejects at run time returnkTfLiteErrorfromInvoke()with a logged message. - Grouped
CONV_2D(input channels a multiple of filter channels) is outside the optimized float kernels' support: FP32 uses the reference kernel, FP16 is rejected atAllocateTensors(). - Pure data-movement operators (
TRANSPOSE,RESHAPE) copy FP16 tensors bitwise, and the FLOAT16 input path ofDEQUANTIZEwidens f16 storage to float32 without f16 arithmetic, so all three work even withoutARM_NN_ENABLE_F16and on Cortex-M4+FP. - HELIA softmax supports only unit beta. FP32 softmax falls back to reference for non-unit beta; FP16 non-unit beta is unsupported.
- HELIA FP16/FP32
UNIDIRECTIONAL_SEQUENCE_LSTMpreserves the TFLite hidden and cell state tensors across invocations when built with ns-cmsis-nn v7.29.0 or newer. With v7.28.0, optimized dispatch still requires zero initial state; FP32 falls back to reference for subsequent invocations, while FP16 stateful execution is unsupported.
See Operator Coverage for the available HELIA operator wrappers.
Non-finite inputs (NaN and infinities)
The optimized floating-point kernels do not all treat NaN the way the TFLM
reference kernels do. The behavior differs by operator, and for TANH it also
differs by target, so it is stated here per case rather than as a single rule.
NaN is not a supported input to the optimized activation kernels
Feeding NaN to the optimized TANH or LOGISTIC kernels does not produce
NaN. It produces a finite value at the activation's saturation bound. If
your model can generate NaN and you rely on it propagating, do not use the
optimized float activation path for that operator.
| Operator | Input | Optimized FP32/FP16 result | TFLM reference result |
|---|---|---|---|
TANH |
NaN, Armv8.1-M MVE targets (Cortex-M55) | Finite, negative, at the saturation bound | NaN |
TANH |
NaN, non-MVE targets (Cortex-M4, Cortex-M3) | NaN | NaN |
TANH |
±Inf | ±1 | ±1 |
LOGISTIC |
NaN, all targets | Finite, at the upper saturation bound (1) | NaN |
LOGISTIC |
+Inf / −Inf | 1 / 0 | 1 / 0 |
ADD, MUL |
NaN | NaN | NaN |
Notes and version boundary:
TANHandLOGISTICare by design. ns-cmsis-nn documents NaN as unsupported input for these kernels: the vectorizedTANHpath usesvminnmq, which is IEEEminNumand returns the numeric operand against a quiet NaN, and theLOGISTICpath clamps its exponent input before evaluation. Restoring NaN would cost a compare and select in the vector loop body.
This did not change in v7.31.0 and is not scheduled to change. ns-cmsis-nn
issue 382 was closed by PR 388, which restored NaN propagation for
RELU/RELU6/LEAKY_RELU/HARDSWISH only -- a different function family.
That PR states explicitly that SIGMOID, TANH and HARDSWISH are outside
the contract it establishes, and it does not touch the TANH or SIGMOID
code paths.
Two details are worth knowing when reading the table above. The scalar
float32 TANH reference carries an explicit if (ax != ax) return x + 0.0f;
NaN guard, which is why the non-MVE row propagates while the MVE row does
not. And there is no MVE LOGISTIC implementation at all for either
float32 or float16 -- sigmoid is always the scalar helper -- which is why the
LOGISTIC row says "all targets" rather than splitting by target like
TANH.
- ADD and MUL are a defect that is now fixed. Up to and including
v7.30.0 the output activation clamp discarded NaN through compare-select
ordering, returning an activation bound instead. ns-cmsis-nn PR 380
reclassifies NaN on the integer bit pattern, which holds at every
optimization level. PR 380 first shipped in v7.31.0, which the heliaRT
pin includes, so ADD and MUL propagate NaN on the optimized path. On
v7.30.0 and earlier they did not.
- The FP32 fallback softens this in practice. Where an operator has a TFLM
reference implementation, HELIA falls back to it when the optimized kernel
declines the configuration, and the reference implementation propagates NaN
normally. FP16 has no reference fallback.
Make builds
The Make integration pins ns-cmsis-nn v7.32.0 and configures the float features
from TARGET_ARCH:
- FP32 is enabled for the HELIA backend.
- FP16 sources and declarations are enabled only for
cortex-m55.
Build an M55 library with FP32 and FP16:
make -f tensorflow/lite/micro/tools/make/Makefile \
TARGET=cortex_m_generic \
TARGET_ARCH=cortex-m55 \
OPTIMIZED_KERNEL_DIR=helia \
microlite
Build an M4+FP library with FP32 only:
make -f tensorflow/lite/micro/tools/make/Makefile \
TARGET=cortex_m_generic \
TARGET_ARCH=cortex-m4+fp \
OPTIMIZED_KERNEL_DIR=helia \
microlite
CMake source builds
NSX entry points
An NSX application does not add the module subdirectories itself. It lists
modules in nsx.yml, and nsx_bootstrap_app() adds them in dependency order —
nsx-cmsis-nn before nsx-helia-rt. The nsx CLI drives CMake, and neither
nsx configure nor nsx build forwards -D options, so the float features
must be set in the application CMakeLists.txt before the bootstrap call:
# ns-cmsis-nn reads these while its own CMakeLists is processed, and they drive
# both its source selection and the ARM_NN_ENABLE_F32/F16 definitions it
# exports. nsx-helia-rt is added afterwards, so it can only observe them.
# Omit either line to build without that float set.
set(ARM_NN_ENABLE_F32 ON CACHE BOOL "" FORCE)
set(ARM_NN_ENABLE_F16 ON CACHE BOOL "" FORCE) # MVEF-capable targets only
include(${CMAKE_CURRENT_LIST_DIR}/cmake/nsx/modules.cmake)
include(${CMAKE_CURRENT_LIST_DIR}/cmake/nsx/nsx_app_bootstrap.cmake)
nsx_bootstrap_app(
APP_ROOT "${CMAKE_CURRENT_LIST_DIR}"
BOARD "${NSX_BOARD}"
MODULES ${NSX_APP_MODULES}
)
target_link_libraries(app PRIVATE nsx::helia_rt)
nsx_finalize_app(app)
Both float sets are opt-in. NSX does not enable either for you, and
omitting the set() lines above is a supported configuration, not an error:
- FP32 off. No
arm_*_f32kernel is linked. FLOAT32 operators still run, on the reference float path, so a float32 model keeps working — it just loses the optimized kernels.nsx-helia-rtprints aNOTICEsaying so. This is what an int8-only image wants: carrying the FP32 kernels cost +31 KB of.text(+11 %) on a KWS DS-CNN with no measurable cycle change (helia-rt#253). - FP16 off. Most helia kernels have no reference float16 path, so a
FLOAT16 operator fails at run time with
kTfLiteErrorand a message naming the type.nsx-helia-rtemits aNOTICE— but only when the target actually has MVE floating point. It is aNOTICErather than aWARNINGbecause the module cannot see whether your model contains a FLOAT16 operator, and int8-only on an MVE-F part is an ordinary configuration.
nsx-helia-rt decides "this target has MVE-F" by compiling a probe with the
active toolchain and board flags (__ARM_FEATURE_MVE & 2), not by matching
CMAKE_SYSTEM_PROCESSOR. A processor name says nothing about whether the
build's -mcpu / -mfpu / +nomve flags left MVE floating point enabled. The
result is cached as HELIA_RT_TARGET_HAS_MVE_FP (an INTERNAL entry, so it
does not appear in cmake -LAH; read it from CMakeCache.txt).
Every helia configure prints one line naming the resolved set and where it came from:
source: is ns-cmsis-nn target when the resolved library itself answered
(the normal case, and the ground truth: it reports what was actually
compiled), or ARM_NN_ENABLE option / Zephyr Kconfig when the target is
not available yet.
Published values
nsx-helia-rt writes the resolved set into the cache so consumers and
generated modules do not have to re-derive it:
| Cache entry | Meaning |
|---|---|
HELIA_RT_FLOAT32_ENABLED |
BOOL. Effective FP32 availability. Output, not a knob. |
HELIA_RT_FLOAT16_ENABLED |
BOOL. Effective FP16 availability. Output, not a knob. |
HELIA_RT_TARGET_HAS_MVE_FP |
INTERNAL. Result of the MVE-F compile probe. Hidden from cmake -LAH; grep CMakeCache.txt for it. |
Both values come from asking the resolved ns-cmsis-nn library what it built:
its float query where the pinned revision exports one, otherwise the compile
definitions on its target. heliaAOT's generated module reads the same query
(helia-aot#386), landed on its
main and shipping in the next heliaAOT release, so the two engines resolve one
answer. heliaRT does not write the
ARM_NN_ENABLE_F32 / ARM_NN_ENABLE_F16 cache entries: those are your
request, and a request the library did not ship is reported as a WARNING.
ns-cmsis-nn revision
ARM_NN_ENABLE_F32/F16 are the only float switches from ns-cmsis-nn
v7.32.0 on: the earlier NSX_CMSIS_NN_ENABLE_F32/F16 spelling was
removed there. An NSX registry that still pins an nsx-cmsis-nn older
than v7.32.0 does not read ARM_NN_ENABLE_* in its NSX module, so
override the module revision in the app's nsx.yml (a module_registry
revision override, or source.path for a local working tree) before
enabling the float features.
Two different GCC 14 fixes, one for each build path
GCC 14 hits an internal compiler error on ns-cmsis-nn's FP16 sources (GCC PR 118460). The two build paths fix it in different places, and neither helps the other:
- Make:
tools/make/ext_libs/helia.incpasses-fno-ssa-phiopt, which works around the ICE with any ns-cmsis-nn revision, because Make compiles the ns-cmsis-nn sources itself. - CMake / NSX / Zephyr: ns-cmsis-nn compiles its own sources, so the workaround is not ours to apply. Use ns-cmsis-nn v7.30.0 or later, which fixed the sources outright.
They are not alternatives to pick between: use whichever belongs to the build path you are on.
Because a static-archive build cannot reveal a missing kernel, verify the final executable rather than the library:
Standalone CMake entry points
When adding the root ns-cmsis-nn project rather than its NSX module, use its standalone option names:
cmake -S app -B build \
-DARM_NN_ENABLE_F32=ON \
-DARM_NN_ENABLE_F16=ON \
-DHELIA_RT_ENABLE_HELIA=ON
cmake --build build
The parent project must add ns-cmsis-nn before heliaRT so
helia_rt::helia can resolve and inspect the dependency target:
add_subdirectory(${NS_CMSIS_NN_DIR} ns-cmsis-nn)
add_subdirectory(${HELIA_RT_DIR} helia-rt)
target_link_libraries(app PRIVATE helia_rt::helia)
With an integer-only ns-cmsis-nn target, generic CMake still builds heliaRT: FP32 uses reference fallbacks, while FP16 operators without reference support remain unavailable.
Source heliaRT with prebuilt ns-cmsis-nn
The ns-cmsis-nn NSX module can expose a prebuilt archive through
NSX_CMSIS_NN_LIB. As with the float features, these cache variables must be
set in the app CMakeLists.txt above nsx_bootstrap_app(), before the
nsx-cmsis-nn module is processed:
set(NSX_CMSIS_NN_LIB "/path/to/libns-cmsis-nn.a" CACHE FILEPATH "" FORCE)
set(NSX_CMSIS_NN_MANIFEST "/path/to/manifest.json" CACHE FILEPATH "" FORCE)
set(ARM_NN_ENABLE_F32 ON CACHE BOOL "" FORCE)
set(ARM_NN_ENABLE_F16 ON CACHE BOOL "" FORCE)
The manifest records which float kernels were compiled into the archive.
ns-cmsis-nn validates the requested features against it and exports matching
ARM_NN_ENABLE_F32/F16 definitions to heliaRT. A prebuilt archive without a
manifest cannot be verified by CMake; ns-cmsis-nn warns and trusts the requested
settings, so use the manifest distributed with the archive.
Zephyr builds
Zephyr resolves Kconfig before module CMake processing, so both modules consume the same final configuration:
CONFIG_HELIA_RT=y
CONFIG_HELIA_RT_BACKEND_HELIA=y
CONFIG_NS_CMSIS_NN_ENABLE_F32=y
CONFIG_NS_CMSIS_NN_ENABLE_F16=y
The HELIA backend implies FP32 and implies FP16 when
ARMV8_1_M_MVEF is available. Explicit settings are useful when auditing a
product configuration. Kconfig prevents FP16 from being selected without MVEF.
Because Kconfig is resolved first, heliaRT and ns-cmsis-nn agree by
construction, and the Zephyr path performs no float-feature validation of its
own. imply is a weak default, though: an explicit
CONFIG_NS_CMSIS_NN_ENABLE_F32=n overrides it and silently drops the HELIA
float operators to reference implementations, with no warning and no error.
Check the generated build/zephyr/.config when auditing a configuration.
Release static libraries
GitHub releases contain combined libhelia-rt-*.a archives. The Make build adds
both heliaRT objects and the selected ns-cmsis-nn kernel objects to the same
archive, so consumers do not link a separate ns-cmsis-nn library.
The release workflow builds:
| Archive target | Included float kernels |
|---|---|
| Cortex-M4+FP | FP32 |
| Cortex-M55 | FP32 and FP16 |
Each architecture is built for GCC, Arm Compiler 6, and ATfE in debug,
release, and release_with_logs variants. The release builder checks the
finished archive for representative FP32 kernel objects and, for M55, FP16
kernel objects. It also rejects FP16 objects in M4+FP archives.
Float support in a prebuilt heliaRT archive is fixed when that archive is created. Application compiler definitions cannot add a kernel omitted from the archive. Always choose the archive matching the target architecture, floating point ABI, toolchain, and build variant.
Migration notes for existing consumers
Enabling the float feature contract changes behavior for integrations built against earlier heliaRT releases:
- Float switch names:
NSX_CMSIS_NN_ENABLE_F32/F16were removed in ns-cmsis-nn v7.32.0 and in this heliaRT release. SetARM_NN_ENABLE_F32/F16instead, in the same place and with the same values. Below ns-cmsis-nn 7.32.0 the old names keep working as before; from 7.32.0 the library rejects them at configure, so rename before you bump the pin. A build directory configured with the old names keeps them inCMakeCache.txt, and the pre-rename NSX module also wrote the new names into the cache, where 7.32.0'soption()keeps whatever value it finds. Clear all six withcmake -U NSX_CMSIS_NN_ENABLE_F32 -U NSX_CMSIS_NN_ENABLE_F16 -U ARM_NN_ENABLE_F32 -U ARM_NN_ENABLE_F16 -U HELIA_RT_ARM_NN_MIRROR_F32 -U HELIA_RT_ARM_NN_MIRROR_F16 <build-dir>. Configuring a fresh build directory is the reliable route, since it leaves no stale cache entry to override the new defaults. The Zephyr symbolsCONFIG_NS_CMSIS_NN_ENABLE_F32/F16are unchanged. - NSX apps: the helia backend no longer requires FP32. An app that never
set
ARM_NN_ENABLE_F32configures and builds int8-only, and pays none of the float code size. Apps that want the optimized float kernels still need theset(... CACHE BOOL "" FORCE)lines shown above beforensx_bootstrap_app(), plus an ns-cmsis-nn of v7.32.0 or newer, the first revision whose NSX module readsARM_NN_ENABLE_F32/F16. That floor also covers the GCC 14 FP16 ICE fixed in v7.30.0 (GCC PR 118460). Releases 1.19.0 and earlier failed configure with aFATAL_ERRORin this situation. - Recovering the size on an int8 app.
ARM_NN_ENABLE_F32/F16default toOFFin ns-cmsis-nn's NSX module from v7.32.0, so nothing enables them for you. If you addedset(NSX_CMSIS_NN_ENABLE_F32 ON CACHE BOOL "" FORCE)only to clear the 1.19.0 configure error, replace it withset(ARM_NN_ENABLE_F32 OFF CACHE BOOL "" FORCE)or drop it: on an int8 model that line is what is costing the ~31 KB. Keep the switchONif you run float32 models and want the optimized kernels. Nothing else in the app needs to change. - Zephyr:
CONFIG_HELIA_RT_BACKEND_HELIAnowimplysNS_CMSIS_NN_ENABLE_F32/F16. If your west workspace pins an ns-cmsis-nn module older than v7.28.0, those Kconfig symbols do not exist and the configuration step emits undefined-symbol warnings; update the module revision to silence them and to get the float kernels. - Library size: Make-based helia builds now always compile the FP32
kernels (and FP16 on
cortex-m55) into the combined archive. Integer-only models still reference them transitively through the operator wrappers, so build applications with-ffunction-sections -fdata-sectionsand link with--gc-sections(the standard embedded configuration) to keep unused float code out of flash.
Troubleshooting
undefined reference to arm_*_f16
: FP16 was enabled in heliaRT declarations but omitted from ns-cmsis-nn. Set
the appropriate CMake option before adding ns-cmsis-nn, or select a prebuilt
archive whose manifest reports FP16 support.
FP16 compiles but faults on target : Confirm that the CPU and build flags enable Armv8.1-M MVE floating point. A Cortex-M55 core can be configured without MVEF.
FP32 unexpectedly uses reference code
: Confirm that the linked ns-cmsis-nn target exports
ARM_NN_ENABLE_F32=1 and that the operator configuration is supported by the
optimized kernel.
Changing CFLAGS has no effect
: Source selection happens during CMake configuration. Set the CMake option
(ARM_NN_ENABLE_F16, the same name on every CMake path) before
ns-cmsis-nn is added. For NSX apps this means a set(... CACHE BOOL "" FORCE)
in the app CMakeLists.txt above nsx_bootstrap_app() — nsx build does not
forward -D options.
nsx-helia-rt: fp32 helia kernels are OFF
: Informational (NOTICE), not an error. The app did not set
ARM_NN_ENABLE_F32 before nsx_bootstrap_app(), or the resolved
nsx-cmsis-nn predates v7.32.0 and does not read it. FLOAT32 operators
run on the reference path. Set the option to get the optimized kernels.
nsx-helia-rt: this target has MVE floating point but the fp16 helia kernels are OFF
: The MVE-F probe succeeded and FP16 is off. FLOAT16 operators will fail at
run time. Set ARM_NN_ENABLE_F16 before nsx_bootstrap_app().
nsx-helia-rt: ARM_NN_ENABLE_F32=... was ignored
: You asked for a float set the linked ns-cmsis-nn target did not compile,
usually because the option was set after ns-cmsis-nn was added, or on a
build directory that had already configured it. heliaRT follows what the
target shipped. Set the option before ns-cmsis-nn is added, in a fresh
build directory.