{
  "$schema": "https://ambiqai.github.io/helia-ui/schema/reference-model-1.json",
  "generatedFrom": {
    "sourceCommit": "5f3fed9f21a57390cc7f00f77a37db8f5f110cb8",
    "tool": "doxyref",
    "version": "1.17.0"
  },
  "language": "c",
  "modules": [
    {
      "description": "Elementwise add and multiplication functions.",
      "name": "Elementwise Functions",
      "path": "heliaCORE.groupElementwise",
      "submodules": [],
      "summary": "Elementwise add and multiplication functions.",
      "symbols": [
        {
          "description": "Elementwise add with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
          "examples": [],
          "id": "arm_elementwise_add_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_add_f32",
          "params": [
            {
              "description": "Pointer to the first input vector.",
              "direction": "in",
              "name": "input_1_vect",
              "type": "const float32_t *"
            },
            {
              "description": "Pointer to the second input vector.",
              "direction": "in",
              "name": "input_2_vect",
              "type": "const float32_t *"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float32_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float32_t"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_add_f32(\n    const float32_t *input_1_vect,\n    const float32_t *input_2_vect,\n    float32_t *output,\n    float32_t out_activation_min,\n    float32_t out_activation_max,\n    int32_t block_size\n)",
          "source": {
            "line": 714,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L714"
          },
          "summary": "Elementwise add with optional output clamp."
        },
        {
          "description": "Elementwise subtract with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
          "examples": [],
          "id": "arm_elementwise_sub_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_sub_f32",
          "params": [
            {
              "description": "Pointer to the first input vector (minuend).",
              "direction": "in",
              "name": "input_1_vect",
              "type": "const float32_t *"
            },
            {
              "description": "Pointer to the second input vector (subtrahend).",
              "direction": "in",
              "name": "input_2_vect",
              "type": "const float32_t *"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float32_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float32_t"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_sub_f32(\n    const float32_t *input_1_vect,\n    const float32_t *input_2_vect,\n    float32_t *output,\n    float32_t out_activation_min,\n    float32_t out_activation_max,\n    int32_t block_size\n)",
          "source": {
            "line": 746,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L746"
          },
          "summary": "Elementwise subtract with optional output clamp."
        },
        {
          "description": "Elementwise absolute value.",
          "examples": [],
          "id": "arm_nn_abs_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_nn_abs_f32",
          "params": [
            {
              "description": "Pointer to the input vector.",
              "direction": "in",
              "name": "input",
              "type": "const float32_t *"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_nn_abs_f32(const float32_t *input, float32_t *output, int32_t block_size)",
          "source": {
            "line": 762,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L762"
          },
          "summary": "Elementwise absolute value."
        },
        {
          "description": "Fill a float32 vector with one value.\n\nBit copy of `value` into every element (vector splat / plain stores), so a NaN fill value lands bit-exact, sign and payload included. Not named arm_fill_f32: CMSIS-DSP exports that symbol.",
          "examples": [],
          "id": "arm_nn_fill_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_nn_fill_f32",
          "params": [
            {
              "description": "Fill value.",
              "direction": "in",
              "name": "value",
              "type": "float32_t"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            },
            {
              "description": "Number of elements to write (0 is a no-op).",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_nn_fill_f32(float32_t value, float32_t *output, int32_t block_size)",
          "source": {
            "line": 778,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L778"
          },
          "summary": "Fill a float32 vector with one value."
        },
        {
          "description": "Elementwise multiply with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
          "examples": [],
          "id": "arm_elementwise_mul_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_mul_f32",
          "params": [
            {
              "description": "Pointer to the first input vector.",
              "direction": "in",
              "name": "input_1_vect",
              "type": "const float32_t *"
            },
            {
              "description": "Pointer to the second input vector.",
              "direction": "in",
              "name": "input_2_vect",
              "type": "const float32_t *"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float32_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float32_t"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_mul_f32(\n    const float32_t *input_1_vect,\n    const float32_t *input_2_vect,\n    float32_t *output,\n    float32_t out_activation_min,\n    float32_t out_activation_max,\n    int32_t block_size\n)",
          "source": {
            "line": 805,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L805"
          },
          "summary": "Elementwise multiply with optional output clamp."
        },
        {
          "description": "Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.\n\nThe result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.",
          "examples": [],
          "id": "arm_minimum_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_minimum_f32",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Input 1, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Input 2, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_minimum_f32(\n    const cmsis_nn_context *ctx,\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
          "source": {
            "line": 858,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L858"
          },
          "summary": "Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape."
        },
        {
          "description": "Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.\n\nThe result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.",
          "examples": [],
          "id": "arm_maximum_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_maximum_f32",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Input 1, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Input 2, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_maximum_f32(\n    const cmsis_nn_context *ctx,\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
          "source": {
            "line": 894,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L894"
          },
          "summary": "Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape."
        },
        {
          "description": "Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and `output_dims` must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.\n\nNumerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged  a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as `scalar - element`, not as the negation of `element - scalar`, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.",
          "examples": [],
          "id": "arm_elementwise_sub_broadcast_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_sub_broadcast_f32",
          "params": [
            {
              "description": "Minuend, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Subtrahend, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float32_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_sub_broadcast_f32(\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t out_activation_min,\n    float32_t out_activation_max\n)",
          "source": {
            "line": 933,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L933"
          },
          "summary": "Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp."
        },
        {
          "description": "Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.",
          "examples": [],
          "id": "arm_elementwise_add_broadcast_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_add_broadcast_f32",
          "params": [
            {
              "description": "First input, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Second input, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float32_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float32_t"
            }
          ],
          "raises": [],
          "returns": [],
          "signature": "arm_cmsis_nn_status arm_elementwise_add_broadcast_f32(\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t out_activation_min,\n    float32_t out_activation_max\n)",
          "source": {
            "line": 957,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L957"
          },
          "summary": "Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp."
        },
        {
          "description": "Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.",
          "examples": [],
          "id": "arm_elementwise_mul_broadcast_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_mul_broadcast_f32",
          "params": [
            {
              "description": "First input, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Second input, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float32_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float32_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float32_t"
            }
          ],
          "raises": [],
          "returns": [],
          "signature": "arm_cmsis_nn_status arm_elementwise_mul_broadcast_f32(\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t out_activation_min,\n    float32_t out_activation_max\n)",
          "source": {
            "line": 981,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L981"
          },
          "summary": "Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp."
        },
        {
          "description": "Elementwise square root.\n\nThe value path is scalar on every toolchain, because Helium has no vector square root. armclang and ATfE do vectorize the surrounding special-value classification; the results are bit-identical to the GCC scalar build, verified by executing both toolchains' objects (#295). Normal positive inputs evaluate `sqrtf(x)`, which IEEE 754 makes correctly rounded, so results are bit-exact to a float64 reference. Subnormal inputs follow FPSCR.FZ: where flush-to-zero is set - the Corstone-300 FVP default, and any host binary linked at -Ofast, where crtfastmath sets DAZ and FTZ - a subnormal input reads as zero and the result is +0. The float16 pair is immune, because it widens to a normal float32 first. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).",
          "examples": [],
          "id": "arm_nn_sqrt_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_nn_sqrt_f32",
          "params": [
            {
              "description": "Pointer to the input vector.",
              "direction": "in",
              "name": "input",
              "type": "const float32_t *"
            },
            {
              "description": "Pointer to the output vector; may alias `input`.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_nn_sqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)",
          "source": {
            "line": 1012,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1012"
          },
          "summary": "Elementwise square root."
        },
        {
          "description": "Elementwise reciprocal square root, `1 / sqrt(x)`.\n\nSame value path as arm_nn_sqrt_f32, including its FPSCR.FZ behaviour on subnormal inputs, where the result is +Inf. Normal positive inputs evaluate `1.0f / sqrtf(x)` in float32: two IEEE roundings, so the result is within 1 ulp of the correctly rounded value (measured against a float64 reference; `x = 4^k` is exact). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).",
          "examples": [],
          "id": "arm_rsqrt_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_rsqrt_f32",
          "params": [
            {
              "description": "Pointer to the input vector.",
              "direction": "in",
              "name": "input",
              "type": "const float32_t *"
            },
            {
              "description": "Pointer to the output vector; may alias `input`.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_rsqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)",
          "source": {
            "line": 1031,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1031"
          },
          "summary": "Elementwise reciprocal square root, 1 / sqrt(x)."
        },
        {
          "description": "Elementwise add with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
          "examples": [],
          "id": "arm_elementwise_add_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_add_f16",
          "params": [
            {
              "description": "Pointer to the first input vector.",
              "direction": "in",
              "name": "input_1_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the second input vector.",
              "direction": "in",
              "name": "input_2_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float16_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float16_t"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_add_f16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    float16_t out_activation_min,\n    float16_t out_activation_max,\n    int32_t block_size\n)",
          "source": {
            "line": 2815,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2815"
          },
          "summary": "Elementwise add with optional output clamp."
        },
        {
          "description": "Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate `arm_elementwise_add_f16()`. New code should call `arm_elementwise_add_f16()` instead.\n\nThis entry does NOT share the contract of `arm_elementwise_add_f16()`:\n\n:::caution\nNo argument validation is performed. A NULL `input_1_vect`, `input_2_vect` or `output` is dereferenced rather than reported. A `block_size` of 0 writes nothing and still returns `ARM_CMSIS_NN_SUCCESS`, where `arm_elementwise_add_f16()` returns `ARM_CMSIS_NN_ARG_ERROR`.\n\n:::\n\n:::note\nThe clamp does not propagate NaN. Both the Helium path (`vminnm`/`vmaxnm`) and the scalar path (the non-propagating `MIN`/`MAX` clamp helper) bound against `out_activation_max` first, so a NaN produced by the addition comes back as `out_activation_max`. `arm_elementwise_add_f16()` documents TensorFlow Lite NaN propagation; this entry does not implement it.\n\n:::",
          "examples": [],
          "id": "arm_elementwise_add_fp16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_add_fp16",
          "params": [
            {
              "description": "Pointer to the first input vector. Must not be NULL.",
              "direction": "in",
              "name": "input_1_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the second input vector. Must not be NULL.",
              "direction": "in",
              "name": "input_2_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the output vector. Must not be NULL.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "const float16_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "const float16_t"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "const int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` unconditionally."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_add_fp16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    const float16_t out_activation_min,\n    const float16_t out_activation_max,\n    const int32_t block_size\n)",
          "source": {
            "line": 2846,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2846"
          },
          "summary": "Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate armelementwiseaddf16()."
        },
        {
          "description": "Elementwise subtract with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
          "examples": [],
          "id": "arm_elementwise_sub_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_sub_f16",
          "params": [
            {
              "description": "Pointer to the first input vector (minuend).",
              "direction": "in",
              "name": "input_1_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the second input vector (subtrahend).",
              "direction": "in",
              "name": "input_2_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float16_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float16_t"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_sub_f16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    float16_t out_activation_min,\n    float16_t out_activation_max,\n    int32_t block_size\n)",
          "source": {
            "line": 2856,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2856"
          },
          "summary": "Elementwise subtract with optional output clamp."
        },
        {
          "description": "Elementwise squared difference of two float16 vectors.\n\nEach output element is calculated as `(input_1_vect[i] - input_2_vect[i])^2`.",
          "examples": [],
          "id": "arm_elementwise_squared_difference_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_squared_difference_f16",
          "params": [
            {
              "description": "Pointer to the first input vector.",
              "direction": "in",
              "name": "input_1_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the second input vector.",
              "direction": "in",
              "name": "input_2_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when an input/output pointer is NULL or `block_size` is less than 1."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_squared_difference_f16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    int32_t block_size\n)",
          "source": {
            "line": 2876,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2876"
          },
          "summary": "Elementwise squared difference of two float16 vectors."
        },
        {
          "description": "Elementwise absolute value.",
          "examples": [],
          "id": "arm_nn_abs_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_nn_abs_f16",
          "params": [
            {
              "description": "Pointer to the input vector.",
              "direction": "in",
              "name": "input",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_nn_abs_f16(const float16_t *input, float16_t *output, int32_t block_size)",
          "source": {
            "line": 2884,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2884"
          },
          "summary": "Elementwise absolute value."
        },
        {
          "description": "Fill a float16 vector with one value; bit copy of `value`, NaN payload included.",
          "examples": [],
          "id": "arm_nn_fill_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_nn_fill_f16",
          "params": [
            {
              "description": "Fill value.",
              "direction": "in",
              "name": "value",
              "type": "float16_t"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Number of elements to write (0 is a no-op).",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_nn_fill_f16(float16_t value, float16_t *output, int32_t block_size)",
          "source": {
            "line": 2896,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2896"
          },
          "summary": "Fill a float16 vector with one value; bit copy of value, NaN payload included."
        },
        {
          "description": "Split a float32 tensor of any rank into several tensors along one axis.\n\nInverse of arm_concatenation_f32; per-split lengths also cover SPLIT_V. Output `s` has the input shape with `input_shape`[axis] replaced by `split_dims`[s]. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input. A dimension of 0 is accepted and copies nothing.",
          "examples": [],
          "id": "arm_split_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_split_f16",
          "params": [
            {
              "description": "Pointer to the flattened (row-major) input.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Number of dimensions in `input_shape` (>= 1).",
              "direction": "in",
              "name": "input_dims",
              "type": "const int32_t"
            },
            {
              "description": "Input shape; `input_shape`[axis] must equal the sum of `split_dims`.",
              "direction": "in",
              "name": "input_shape",
              "type": "const int32_t *"
            },
            {
              "description": "Axis to split along (0 <= axis < input_dims).",
              "direction": "in",
              "name": "axis",
              "type": "const int32_t"
            },
            {
              "description": "Number of outputs (>= 1).",
              "direction": "in",
              "name": "num_splits",
              "type": "const int32_t"
            },
            {
              "description": "Array of length `num_splits:` each output's extent along `axis`.",
              "direction": "in",
              "name": "split_dims",
              "type": "const int32_t *"
            },
            {
              "description": "Array of `num_splits` pointers to the flattened outputs.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *const *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, split entry, split sum, NULL pointer or an element count above INT32_MAX."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_split_f16(\n    const float16_t *input_data,\n    const int32_t input_dims,\n    const int32_t *input_shape,\n    const int32_t axis,\n    const int32_t num_splits,\n    const int32_t *split_dims,\n    float16_t *const *output_data\n)",
          "source": {
            "line": 2923,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2923"
          },
          "summary": "Split a float32 tensor of any rank into several tensors along one axis."
        },
        {
          "description": "Strided slice for float32 data (pure copy, TensorFlow Lite compatible).",
          "examples": [],
          "id": "arm_strided_slice_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_strided_slice_f16",
          "params": [
            {
              "description": "Pointer to input tensor.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to output tensor.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *const"
            },
            {
              "description": "Begin dimensions for slicing.",
              "direction": "in",
              "name": "begin_dims",
              "type": "const cmsis_nn_dims *const"
            },
            {
              "description": "Stride dimensions for slicing.",
              "direction": "in",
              "name": "stride_dims",
              "type": "const cmsis_nn_dims *const"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *const"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "ARM_CMSIS_NN_SUCCESS on success."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_strided_slice_f16(\n    const float16_t *input_data,\n    float16_t *output_data,\n    const cmsis_nn_dims *const input_dims,\n    const cmsis_nn_dims *const begin_dims,\n    const cmsis_nn_dims *const stride_dims,\n    const cmsis_nn_dims *const output_dims\n)",
          "source": {
            "line": 2934,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2934"
          },
          "summary": "Strided slice for float32 data (pure copy, TensorFlow Lite compatible)."
        },
        {
          "description": "Elementwise multiply with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
          "examples": [],
          "id": "arm_elementwise_mul_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_mul_f16",
          "params": [
            {
              "description": "Pointer to the first input vector.",
              "direction": "in",
              "name": "input_1_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the second input vector.",
              "direction": "in",
              "name": "input_2_vect",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the output vector.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float16_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float16_t"
            },
            {
              "description": "Number of elements to process.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_mul_f16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    float16_t out_activation_min,\n    float16_t out_activation_max,\n    int32_t block_size\n)",
          "source": {
            "line": 2944,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2944"
          },
          "summary": "Elementwise multiply with optional output clamp."
        },
        {
          "description": "Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.\n\nThe result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.",
          "examples": [],
          "id": "arm_minimum_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_minimum_f16",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Input 1, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Input 2, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_minimum_f16(\n    const cmsis_nn_context *ctx,\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
          "source": {
            "line": 2954,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2954"
          },
          "summary": "Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape."
        },
        {
          "description": "Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.\n\nThe result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.",
          "examples": [],
          "id": "arm_maximum_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_maximum_f16",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Input 1, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Input 2, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_maximum_f16(\n    const cmsis_nn_context *ctx,\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
          "source": {
            "line": 2965,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2965"
          },
          "summary": "Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape."
        },
        {
          "description": "Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and `output_dims` must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.\n\nNumerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged  a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as `scalar - element`, not as the negation of `element - scalar`, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.\n\nHalf-precision twin: the numerics are those of arm_elementwise_sub_f16 on the materialised operands.",
          "examples": [],
          "id": "arm_elementwise_sub_broadcast_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_sub_broadcast_f16",
          "params": [
            {
              "description": "Minuend, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Subtrahend, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float16_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float16_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_elementwise_sub_broadcast_f16(\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t out_activation_min,\n    float16_t out_activation_max\n)",
          "source": {
            "line": 2978,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2978"
          },
          "summary": "Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp."
        },
        {
          "description": "Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.\n\nHalf-precision twin: the numerics are those of arm_elementwise_add_f16 on the materialised operands.",
          "examples": [],
          "id": "arm_elementwise_add_broadcast_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_add_broadcast_f16",
          "params": [
            {
              "description": "First input, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Second input, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float16_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float16_t"
            }
          ],
          "raises": [],
          "returns": [],
          "signature": "arm_cmsis_nn_status arm_elementwise_add_broadcast_f16(\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t out_activation_min,\n    float16_t out_activation_max\n)",
          "source": {
            "line": 2992,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2992"
          },
          "summary": "Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp."
        },
        {
          "description": "Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.\n\nHalf-precision twin: the numerics are those of arm_elementwise_mul_f16 on the materialised operands.",
          "examples": [],
          "id": "arm_elementwise_mul_broadcast_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_elementwise_mul_broadcast_f16",
          "params": [
            {
              "description": "First input, NHWC, sized by `input_1_dims`.",
              "direction": "in",
              "name": "input_1_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 1.",
              "direction": "in",
              "name": "input_1_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Second input, NHWC, sized by `input_2_dims`.",
              "direction": "in",
              "name": "input_2_data",
              "type": "const float16_t *"
            },
            {
              "description": "Dimensions of input 2.",
              "direction": "in",
              "name": "input_2_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output, NHWC, sized by `output_dims`.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Broadcast output dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Minimum output clamp value.",
              "direction": "in",
              "name": "out_activation_min",
              "type": "float16_t"
            },
            {
              "description": "Maximum output clamp value.",
              "direction": "in",
              "name": "out_activation_max",
              "type": "float16_t"
            }
          ],
          "raises": [],
          "returns": [],
          "signature": "arm_cmsis_nn_status arm_elementwise_mul_broadcast_f16(\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t out_activation_min,\n    float16_t out_activation_max\n)",
          "source": {
            "line": 3006,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3006"
          },
          "summary": "Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp."
        },
        {
          "description": "Elementwise square root of a float16 tensor.\n\nThe value path is scalar on every toolchain, because Helium has no vector square root; armclang and ATfE vectorize the surrounding classification into an MVE loop and produce bit-identical results, verified by executing their objects (#295). Each element is widened to float32, `sqrtf` is evaluated there and the result is rounded once to float16. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 square root (0 ulp, #295). Widening first also makes this pair immune to FPSCR.FZ, which flushes float32 subnormals in the f32 pair. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).",
          "examples": [],
          "id": "arm_nn_sqrt_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_nn_sqrt_f16",
          "params": [
            {
              "description": "Pointer to the input tensor.",
              "direction": "in",
              "name": "input",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the output tensor; may alias `input`.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Number of tensor elements.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_nn_sqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)",
          "source": {
            "line": 3036,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3036"
          },
          "summary": "Elementwise square root of a float16 tensor."
        },
        {
          "description": "Elementwise reciprocal square root of a float16 tensor, `1 / sqrt(x)`.\n\nSame value path as arm_nn_sqrt_f16: widen to float32, evaluate `1.0f / sqrtf(x)` there, round once to float16, and so also immune to FPSCR.FZ. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 reciprocal square root (0 ulp, #295). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).",
          "examples": [],
          "id": "arm_rsqrt_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_rsqrt_f16",
          "params": [
            {
              "description": "Pointer to the input tensor.",
              "direction": "in",
              "name": "input",
              "type": "const float16_t *"
            },
            {
              "description": "Pointer to the output tensor; may alias `input`.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Number of tensor elements.",
              "direction": "in",
              "name": "block_size",
              "type": "int32_t"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_rsqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)",
          "source": {
            "line": 3055,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3055"
          },
          "summary": "Elementwise reciprocal square root of a float16 tensor, 1 / sqrt(x)."
        }
      ]
    }
  ],
  "name": "heliaCORE"
}
