{
  "$schema": "https://ambiqai.github.io/helia-ui/schema/reference-model-1.json",
  "generatedFrom": {
    "sourceCommit": "5f3fed9f21a57390cc7f00f77a37db8f5f110cb8",
    "tool": "doxyref",
    "version": "1.17.0"
  },
  "language": "c",
  "modules": [
    {
      "description": "",
      "name": "heliaCORE",
      "path": "heliaCORE",
      "submodules": [
        {
          "description": "Perform activation layers, including ReLU (Rectified Linear Unit), sigmoid and tanh",
          "name": "Activation Functions",
          "path": "heliaCORE.Acti",
          "submodules": [],
          "summary": "Perform activation layers, including ReLU (Rectified Linear Unit), sigmoid and tanh",
          "symbols": [
            {
              "description": "Elementwise activation.\n\n:::note\nThe RELU, RELU6 and LEAKY_RELU legs classify NaN on the integer bit pattern (#380 / #382), so a NaN input comes back as NaN at every optimization level on the gated toolchains, including the shipped -Ofast. This holds on both the scalar and the MVE (cortex-m55) build paths; the MVE RELU/RELU6 legs restore the NaN lanes that vmaxnmq/vminnmq suppress. The MVE TANH leg returns a NaN input unchanged and keeps the sign of zero, also decided on the bit pattern (#635). The scalar TANH leg returns NaN where there is no hardware floating point (__ARM_FP undefined, e.g. Cortex-M0), where it too classifies NaN on the bit pattern (quieting a signalling NaN), and elsewhere only in builds without -ffinite-math-only. SIGMOID and HARDSWISH are outside this contract; see the per-helper notes in `Include/Internal/arm_nn_activation_flt.h`.\n\n:::\n\n:::note\nThe HARDSWISH leg's scalar helper (arm_nn_hardswish_scalar_f32, serving every build that does not take the MVE float path  no MVE float support, or MVE present but not used, e.g. under ARM_MATH_AUTOVECTORIZE) keeps the legacy separately rounded multiply-and-add gate and can differ by an ulp in the curved region from the standalone `arm_hard_swish_f32`, whose gate is a correctly rounded fma; the mux's MVE helper (arm_nn_vhardswish_mve_f32) uses vfmaq and agrees with that kernel. Callers that need bit-exact, leg-agreeing hard swish  or the documented NaN/Inf contract  should call `arm_hard_swish_f32` directly.\n\n:::",
              "examples": [],
              "id": "arm_nn_activation_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_activation_f32",
              "params": [
                {
                  "description": "Pointer to the input samples.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output samples.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                },
                {
                  "description": "Activation selector.",
                  "direction": "in",
                  "name": "type",
                  "type": "arm_nn_activation_type_flt"
                },
                {
                  "description": "Extra activation parameter. Used for parameterized activations such as leaky ReLU.",
                  "direction": "in",
                  "name": "act_param",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_activation_f32(\n    const float32_t *input,\n    float32_t *output,\n    int32_t size,\n    arm_nn_activation_type_flt type,\n    float32_t act_param\n)",
              "source": {
                "line": 609,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L609"
              },
              "summary": "Elementwise activation."
            },
            {
              "description": "Parametric ReLU for float32 data.\n\nComputes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.",
              "examples": [],
              "id": "arm_prelu_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_prelu_f32",
              "params": [
                {
                  "description": "Input tensor dimensions. Must equal output_dims.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Alpha tensor dimensions.",
                  "direction": "in",
                  "name": "alpha_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the alpha (slope) tensor.",
                  "direction": "in",
                  "name": "alpha",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_prelu_f32(\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *alpha_dims,\n    const float32_t *alpha,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output\n)",
              "source": {
                "line": 631,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L631"
              },
              "summary": "Parametric ReLU for float32 data."
            },
            {
              "description": "Hard swish activation for float32 data.\n\nComputes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise, evaluated as x * clamp(fma(x, 1/6, 0.5), 0, 1) so the saturated regions are exact: x >= 3 returns x bit-exactly and x <= -3 returns zero exactly (a negative zero, as IEEE negative * +0.0). In the curved region -3 < x < 3 the gate is a correctly rounded fused multiply-add on both build paths, so the scalar and MVE (cortex-m55) legs agree bit-exactly on every numeric normal input. Two carve-outs, both rooted in Armv8.1-M MVE floating-point arithmetic using the architecture's Standard FPSCR value  DN=1 and FZ=1 hard-wired, FZ16 passed through (Arm v8-M ARM, DDI 0553B.l, StandardFPSCRValue(), selected by the MVE FP pseudocode's fpscr_controlled=FALSE): NaN lanes agree in NaN-ness but not necessarily in payload (forced DN makes the MVE leg canonicalize payloads the scalar leg preserves), and the MVE leg flushes f32 subnormal operands and results to a signed zero regardless of FPSCR.FZ, where the scalar leg with FZ clear keeps them. Both reference models (FVP Corstone-300 and QEMU mps3-an547) exhibit the flush identically; it has not been executed on silicon, where the same architectural behavior is required. Near the lower knot the absolute contract is the meaningful one: for x just above -3 the output error is dominated by the gate constant's representation error, bounded by |x^2 * (1/6f - 1/6)| ~ 4.5e-08 near x = -3 (e.g. nextafterf(-3, 0) returns -7.45e-08 against a float64 -1.19e-07  millions of ulps of the tiny result, well inside the 1e-6 absolute contract), and where the gate underflows to exactly zero the kernel returns -0.0 with unbounded relative error. In-place operation (output == input) is supported on both legs; each element is read before it is written.\n\n:::note\nNaN propagates (TensorFlow Lite semantics): a NaN input element yields NaN at that output element at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast: propagation rides the final multiply x * gate  a NaN x makes the product NaN whatever the gate resolved to  rather than a compare-and-select that -ffinite-math-only could fold. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. +Inf returns +Inf (the gate is 1). -Inf returns NaN, not the mathematical limit 0: the gate is 0 there and (-Inf) * 0 is NaN by IEEE 754, the same result TFLite's float hard-swish reference produces; special-casing -Inf would put a per-element select in the hot loop for an input no finite model produces. The scalar and MVE legs agree on the NaN-ness and on +/-Inf; NaN payload bits may differ between legs.\n\n:::",
              "examples": [],
              "id": "arm_hard_swish_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_hard_swish_f32",
              "params": [
                {
                  "description": "Pointer to the input samples.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output samples.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of elements to process. Must be at least 1.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_hard_swish_f32(const float32_t *input, float32_t *output, int32_t size)",
              "source": {
                "line": 680,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L680"
              },
              "summary": "Hard swish activation for float32 data."
            },
            {
              "description": "Elementwise activation.\n\n:::note\nThe RELU, RELU6 and LEAKY_RELU legs classify NaN on the integer bit pattern (#380 / #382), so a NaN input comes back as NaN at every optimization level on the gated toolchains, including the shipped -Ofast. This holds on both the scalar and the MVE (cortex-m55) build paths; the MVE RELU/RELU6 legs restore the NaN lanes that vmaxnmq/vminnmq suppress. The MVE TANH leg returns a NaN input unchanged and keeps the sign of zero, also decided on the bit pattern (#635). The scalar TANH leg returns NaN where there is no hardware floating point (__ARM_FP undefined, e.g. Cortex-M0), where it too classifies NaN on the bit pattern (quieting a signalling NaN), and elsewhere only in builds without -ffinite-math-only. SIGMOID and HARDSWISH are outside this contract; see the per-helper notes in `Include/Internal/arm_nn_activation_flt.h`.\n\n:::\n\n:::note\nThe HARDSWISH leg's scalar helper (arm_nn_hardswish_scalar_f32, serving every build that does not take the MVE float path  no MVE float support, or MVE present but not used, e.g. under ARM_MATH_AUTOVECTORIZE) keeps the legacy separately rounded multiply-and-add gate and can differ by an ulp in the curved region from the standalone `arm_hard_swish_f32`, whose gate is a correctly rounded fma; the mux's MVE helper (arm_nn_vhardswish_mve_f32) uses vfmaq and agrees with that kernel. Callers that need bit-exact, leg-agreeing hard swish  or the documented NaN/Inf contract  should call `arm_hard_swish_f32` directly.\n\n:::\n\n:::note\nThe RELU, RELU6 and LEAKY_RELU legs classify NaN on the integer bit pattern (#380 / #382), so a NaN input comes back as NaN at every optimization level on the gated toolchains, including the shipped -Ofast. This holds uniformly across build paths: the scalar path serves every build without MVE float16 (and LEAKY_RELU on MVE builds too), while the MVE RELU/RELU6 legs (cortex-m55) restore the NaN lanes that vmaxnmq/vminnmq suppress, using the same integer-domain lane classification as the elementwise clamps. SIGMOID and HARDSWISH are outside this contract; see the per-helper notes in `Include/Internal/arm_nn_activation_flt.h`.\n\n:::\n\n:::note\nTANH propagates NaN on both scalar and MVE paths, preserves the sign of zero, and maps +/-Inf to +/-1, including under -Ofast. NaN payload, sign and signaling state are not specified. The finite LUT interpolation may round differently across paths; bitwise scalar/MVE agreement is not required. Caller FP control settings are not changed.\n\n:::\n\n:::note\nBoth legs of the HARDSWISH mux evaluate natively in float16  the scalar helper (arm_nn_hardswish_scalar_f16) with a separately rounded multiply-and-add gate, the MVE helper (arm_nn_vhardswish_mve_f16) with a float16 vfmaq  so either can differ by an ulp from the scalar leg of the standalone `arm_hard_swish_f16`, which computes in float32 with an fma gate and rounds to float16 once. Callers that need the documented NaN/Inf contract should call `arm_hard_swish_f16` directly.\n\n:::",
              "examples": [],
              "id": "arm_nn_activation_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_activation_f16",
              "params": [
                {
                  "description": "Pointer to the input samples.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output samples.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                },
                {
                  "description": "Activation selector.",
                  "direction": "in",
                  "name": "type",
                  "type": "arm_nn_activation_type_flt"
                },
                {
                  "description": "Extra activation parameter. Used for parameterized activations such as leaky ReLU.",
                  "direction": "in",
                  "name": "act_param",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_activation_f16(\n    const float16_t *input,\n    float16_t *output,\n    int32_t size,\n    arm_nn_activation_type_flt type,\n    float16_t act_param\n)",
              "source": {
                "line": 2750,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2750"
              },
              "summary": "Elementwise activation."
            },
            {
              "description": "Parametric ReLU for float32 data.\n\nComputes output = input >= 0 ? input : input * alpha, with alpha broadcast onto the input (TensorFlow Lite semantics): each alpha dimension must equal the matching input dimension or 1.",
              "examples": [],
              "id": "arm_prelu_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_prelu_f16",
              "params": [
                {
                  "description": "Input tensor dimensions. Must equal output_dims.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Alpha tensor dimensions.",
                  "direction": "in",
                  "name": "alpha_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the alpha (slope) tensor.",
                  "direction": "in",
                  "name": "alpha",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_prelu_f16(\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *alpha_dims,\n    const float16_t *alpha,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
              "source": {
                "line": 2759,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2759"
              },
              "summary": "Parametric ReLU for float32 data."
            },
            {
              "description": "Hard swish activation for float16 data.\n\nComputes output[i] = input[i] * min(max(input[i] + 3, 0), 6) / 6 elementwise. The scalar leg widens each element to float32, evaluates the gate and the product there exactly as in `arm_hard_swish_f32`, and narrows only the final product, so it is single-rounded. The MVE (cortex-m55) leg evaluates the same expression in float16 throughout, scaling the gate by 1/6 before the product so that the multiplier stays in [0, 1]; it rounds the gate and the product separately and so can sit up to 2 float16 ulp away from the scalar leg in the curved region -3 < x < 3. The saturated regions are exact and identical on both legs (x >= 3 returns x bit-exactly, x <= -3 returns zero), as is the NaN/Inf behavior below; NaN lanes agree in NaN-ness but not necessarily in payload. In-place operation (output == input) is supported on both legs.\n\n:::note\nNaN and Inf behave as in `arm_hard_swish_f32`, at every optimization level on the gated toolchains (see docs/guides/verification.md): NaN propagates through the final multiply (NaN-ness only, not a particular payload), +Inf returns +Inf, and -Inf returns NaN because the gate is 0 there and (-Inf) * 0 is NaN by IEEE 754, matching TFLite's float hard-swish reference rather than the mathematical limit 0.\n\n:::\n\n:::note\nNothing in this kernel converts between half and single precision any more, and the float16 kernels that still do are not tied to a particular assembler: they go through `Include/Internal/arm_nn_vcvt_f16.h`, which emits the scalar form of VCVTB/VCVTT wherever the vector form would be mis-encoded (binutils below 2.43). Under CMake the probe measures the assembler in use and selects the form; a build that never runs it  the CMSIS-Pack `Source` Cvariant, `module.mk`, or a CMake project that wires its architecture flags where the probe cannot read them  falls back to the compiler major, which is right for every Arm GNU release and wrong only for a GCC 14 or newer driver paired by hand with an older binutils. Check `as --version` if you assembled that pair yourself. See docs/guides/toolchains.md.\n\n:::",
              "examples": [],
              "id": "arm_hard_swish_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_hard_swish_f16",
              "params": [
                {
                  "description": "Pointer to the input samples.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output samples.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of elements to process. Must be at least 1.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_hard_swish_f16(const float16_t *input, float16_t *output, int32_t size)",
              "source": {
                "line": 2803,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2803"
              },
              "summary": "Hard swish activation for float16 data."
            }
          ]
        },
        {
          "description": "Comparison operators with optional broadcasting support.",
          "name": "Comparison Functions",
          "path": "heliaCORE.Comparison",
          "submodules": [],
          "summary": "Comparison operators with optional broadcasting support.",
          "symbols": []
        },
        {
          "description": "Collection of fully-connected and matrix multiplication functions.\n\nFully-connected layer is basically a matrix-vector multiplication with bias. The matrix is the weights and the input/output vectors are the activation values. Supported {weight, activation} precisions include {8-bit, 8-bit} and {8-bit, 16-bit}",
          "name": "Fully-connected Layer Functions",
          "path": "heliaCORE.FC",
          "submodules": [],
          "summary": "Collection of fully-connected and matrix multiplication functions.",
          "symbols": [
            {
              "description": "Fully connected layer, NHWC layout.",
              "examples": [],
              "id": "arm_fully_connected_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_nhwc_f32",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully connected parameters and activation clamp.",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_nhwc_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params_f32 *fc_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output\n)",
              "source": {
                "line": 1056,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1056"
              },
              "summary": "Fully connected layer, NHWC layout."
            },
            {
              "description": "Fully connected layer, dispatch by layout.",
              "examples": [],
              "id": "arm_fully_connected_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_f32",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully connected parameters and activation clamp.",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params_f32 *fc_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 1084,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1084"
              },
              "summary": "Fully connected layer, dispatch by layout."
            },
            {
              "description": "Get the temporary buffer size required by the fully connected layer.",
              "examples": [],
              "id": "arm_fully_connected_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_f32_get_buffer_size",
              "params": [
                {
                  "description": "Fully connected parameters.",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_fully_connected_f32_get_buffer_size(\n    const cmsis_nn_fc_params_f32 *fc_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 1107,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1107"
              },
              "summary": "Get the temporary buffer size required by the fully connected layer."
            },
            {
              "description": "Batched matrix multiplication.",
              "examples": [],
              "id": "arm_batch_matmul_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_matmul_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Batch matmul parameters and activation clamp.",
                  "direction": "in",
                  "name": "bmm_params",
                  "type": "const cmsis_nn_bmm_params_f32 *"
                },
                {
                  "description": "Left-hand-side input tensor dimensions.",
                  "direction": "in",
                  "name": "input_lhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the left-hand-side input tensor.",
                  "direction": "in",
                  "name": "input_lhs",
                  "type": "const float32_t *"
                },
                {
                  "description": "Right-hand-side input tensor dimensions.",
                  "direction": "in",
                  "name": "input_rhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the right-hand-side input tensor. With `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` each `[K, N]` matrix occupies `K * ceil(N / block) * block` elements (block is 4 for float32, 8 for float16) and consecutive batch matrices are stored back to back at that stride.",
                  "direction": "in",
                  "name": "input_rhs",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_batch_matmul_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_bmm_params_f32 *bmm_params,\n    const cmsis_nn_dims *input_lhs_dims,\n    const float32_t *input_lhs,\n    const cmsis_nn_dims *input_rhs_dims,\n    const float32_t *input_rhs,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output\n)",
              "source": {
                "line": 1491,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1491"
              },
              "summary": "Batched matrix multiplication."
            },
            {
              "description": "Get the temporary buffer size required by batched matrix multiplication.",
              "examples": [],
              "id": "arm_batch_matmul_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_matmul_f32_get_buffer_size",
              "params": [
                {
                  "description": "Batch matmul parameters.",
                  "direction": "in",
                  "name": "bmm_params",
                  "type": "const cmsis_nn_bmm_params_f32 *"
                },
                {
                  "description": "Left-hand-side input tensor dimensions.",
                  "direction": "in",
                  "name": "input_lhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Right-hand-side input tensor dimensions.",
                  "direction": "in",
                  "name": "input_rhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_batch_matmul_f32_get_buffer_size(\n    const cmsis_nn_bmm_params_f32 *bmm_params,\n    const cmsis_nn_dims *input_lhs_dims,\n    const cmsis_nn_dims *input_rhs_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1510,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1510"
              },
              "summary": "Get the temporary buffer size required by batched matrix multiplication."
            },
            {
              "description": "Fully connected layer, NHWC layout.",
              "examples": [],
              "id": "arm_fully_connected_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_nhwc_f16",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully connected parameters and activation clamp.",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_nhwc_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params_f16 *fc_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
              "source": {
                "line": 3067,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3067"
              },
              "summary": "Fully connected layer, NHWC layout."
            },
            {
              "description": "Fully connected layer, NHWC layout.\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_fully_connected_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_fully_connected_nhwc_f16.\n\n:::",
              "examples": [],
              "id": "arm_fully_connected_nhwc_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_nhwc_f16_acc16",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully connected parameters and activation clamp.",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_nhwc_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params_f16 *fc_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
              "source": {
                "line": 3086,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3086"
              },
              "summary": "Fully connected layer, NHWC layout."
            },
            {
              "description": "Fully connected layer, dispatch by layout.\n\n:::note\nAccumulation width follows the matmul helper the weight format selects (arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16): the scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457); the MVE legs use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout.\n\n:::",
              "examples": [],
              "id": "arm_fully_connected_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_f16",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully connected parameters and activation clamp.",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params_f16 *fc_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 3110,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3110"
              },
              "summary": "Fully connected layer, dispatch by layout."
            },
            {
              "description": "Fully connected layer, dispatch by layout.\n\n:::note\nAccumulation width follows the matmul helper the weight format selects (arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16): the scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457); the MVE legs use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_fully_connected_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_fully_connected_f16.\n\n:::",
              "examples": [],
              "id": "arm_fully_connected_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_f16_acc16",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully connected parameters and activation clamp.",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params_f16 *fc_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 3130,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3130"
              },
              "summary": "Fully connected layer, dispatch by layout."
            },
            {
              "description": "Get the temporary buffer size required by the fully connected layer.",
              "examples": [],
              "id": "arm_fully_connected_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_f16_get_buffer_size",
              "params": [
                {
                  "description": "Fully connected parameters.",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_fully_connected_f16_get_buffer_size(\n    const cmsis_nn_fc_params_f16 *fc_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 3145,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3145"
              },
              "summary": "Get the temporary buffer size required by the fully connected layer."
            },
            {
              "description": "Batched matrix multiplication.\n\n:::note\nAccumulation width. Without adjoints the product goes through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16 and so takes their rule: on the MVE legs a reduction of more than 32 taps per output accumulates blockwise (AmbiqAI/ns-cmsis-nn#586), in float16 up to 32. The adjoint paths accumulate in float16 throughout. There is no `_acc16` entry; a caller that needs float16 lanes on a long reduction calls arm_nn_mat_mult_nt_t_f16_acc16 / arm_nn_mat_mult_nt_n_packed_f16_acc16 per batch.\n\n:::",
              "examples": [],
              "id": "arm_batch_matmul_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_matmul_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Batch matmul parameters and activation clamp.",
                  "direction": "in",
                  "name": "bmm_params",
                  "type": "const cmsis_nn_bmm_params_f16 *"
                },
                {
                  "description": "Left-hand-side input tensor dimensions.",
                  "direction": "in",
                  "name": "input_lhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the left-hand-side input tensor.",
                  "direction": "in",
                  "name": "input_lhs",
                  "type": "const float16_t *"
                },
                {
                  "description": "Right-hand-side input tensor dimensions.",
                  "direction": "in",
                  "name": "input_rhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the right-hand-side input tensor. With `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` each `[K, N]` matrix occupies `K * ceil(N / block) * block` elements (block is 4 for float32, 8 for float16) and consecutive batch matrices are stored back to back at that stride.",
                  "direction": "in",
                  "name": "input_rhs",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_batch_matmul_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_bmm_params_f16 *bmm_params,\n    const cmsis_nn_dims *input_lhs_dims,\n    const float16_t *input_lhs,\n    const cmsis_nn_dims *input_rhs_dims,\n    const float16_t *input_rhs,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
              "source": {
                "line": 3327,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3327"
              },
              "summary": "Batched matrix multiplication."
            },
            {
              "description": "Get the temporary buffer size required by batched matrix multiplication.",
              "examples": [],
              "id": "arm_batch_matmul_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_matmul_f16_get_buffer_size",
              "params": [
                {
                  "description": "Batch matmul parameters.",
                  "direction": "in",
                  "name": "bmm_params",
                  "type": "const cmsis_nn_bmm_params_f16 *"
                },
                {
                  "description": "Left-hand-side input tensor dimensions.",
                  "direction": "in",
                  "name": "input_lhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Right-hand-side input tensor dimensions.",
                  "direction": "in",
                  "name": "input_rhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_batch_matmul_f16_get_buffer_size(\n    const cmsis_nn_bmm_params_f16 *bmm_params,\n    const cmsis_nn_dims *input_lhs_dims,\n    const cmsis_nn_dims *input_rhs_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 3339,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3339"
              },
              "summary": "Get the temporary buffer size required by batched matrix multiplication."
            }
          ]
        },
        {
          "description": "",
          "name": "Gather Functions:",
          "path": "heliaCORE.Gather",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Gather contiguous slices along an axis.\n\nData rank is 1..4 and indices rank is 0..4; rank-0 indices contain one index. Negative axis normalizes by input_rank; negative batch_dims normalizes by coords_rank. After normalization, 0 <= batch_dims <= coords_rank and batch_dims <= axis < input_rank. Leading batch dimensions must match. The inferred output shape is input_shape[:axis] + indices_shape[batch_dims:] + input_shape[axis + 1:].\n\nShapes use the first rank fields of `cmsis_nn_dims` in n, h, w, c order; unused fields are ignored. The inferred output rank must be 0..4, and output_dims must match its leading dimensions. A rank-0 output is one element. All dimension extents must be nonnegative. Input, index and output buffer byte counts must each fit INT32_MAX; this is a CORE capacity limit.\n\nAll metadata pointers are required. A NULL data, indices or output pointer is accepted only when that respective buffer has zero elements. Valid empty calls copy nothing. All supplied indices are checked, even for empty output. Coordinates must be nonnegative and below their corresponding axis extent. Invalid metadata or indices return ARG_ERROR without changing output.\n\nThis operation preserves all bits, including NaN payloads, signed zero and subnormals, independently of floating-point controls. Buffers must not overlap. No scratch buffer is required. Portable copies use existing MVE copy paths when enabled.",
              "examples": [],
              "id": "arm_gather_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gather_f32",
              "params": [
                {
                  "description": "Input data buffer.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Input shape in leading-dimension order.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Signed 32-bit indices.",
                  "direction": "in",
                  "name": "indices_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Indices shape in leading-dimension order.",
                  "direction": "in",
                  "name": "indices_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Ranks and gathering parameters.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gather_params *"
                },
                {
                  "description": "Output data buffer.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Inferred output shape in leading-dimension order.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gather_f32(\n    const float32_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t *indices_data,\n    const cmsis_nn_dims *indices_dims,\n    const cmsis_nn_gather_params *params,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2135,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2135"
              },
              "summary": "Gather contiguous slices along an axis."
            },
            {
              "description": "Gather contiguous slices using coordinate tuples.\n\nData rank is 1..4 and indices rank is 1..4. The final indices dimension is the tuple width, which must be at least one. batch_dims is a TensorFlow-style extension (not a LiteRT builtin option): 0 <= batch_dims < indices_rank, batch_dims < params_rank, and batch_dims + tuple_width <= params_rank. Leading batch dimensions must match. The inferred output shape is indices_shape[:-1] + params_shape[batch_dims + tuple_width:]. Empty data with a nonempty index buffer is rejected.\n\nShapes use the first rank fields of `cmsis_nn_dims` in n, h, w, c order; unused fields are ignored. The inferred output rank must be 0..4, and output_dims must match its leading dimensions. A rank-0 output is one element. All dimension extents must be nonnegative. Input, index and output buffer byte counts must each fit INT32_MAX; this is a CORE capacity limit.\n\nAll metadata pointers are required. A NULL data, indices or output pointer is accepted only when that respective buffer has zero elements. Valid empty calls copy nothing. All supplied indices are checked, even for empty output. Coordinates must be nonnegative and below their corresponding axis extent. Invalid metadata or indices return ARG_ERROR without changing output.\n\nThis operation preserves all bits, including NaN payloads, signed zero and subnormals, independently of floating-point controls. Buffers must not overlap. No scratch buffer is required. Portable copies use existing MVE copy paths when enabled.",
              "examples": [],
              "id": "arm_gather_nd_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gather_nd_f32",
              "params": [
                {
                  "description": "Input data buffer.",
                  "direction": "in",
                  "name": "params_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Input shape in leading-dimension order.",
                  "direction": "in",
                  "name": "params_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Signed 32-bit indices.",
                  "direction": "in",
                  "name": "indices_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Indices shape in leading-dimension order.",
                  "direction": "in",
                  "name": "indices_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Ranks and gathering parameters.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gather_nd_params *"
                },
                {
                  "description": "Output data buffer.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Inferred output shape in leading-dimension order.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gather_nd_f32(\n    const float32_t *params_data,\n    const cmsis_nn_dims *params_dims,\n    const int32_t *indices_data,\n    const cmsis_nn_dims *indices_dims,\n    const cmsis_nn_gather_nd_params *params,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2180,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2180"
              },
              "summary": "Gather contiguous slices using coordinate tuples."
            },
            {
              "description": "Gather contiguous slices along an axis.\n\nData rank is 1..4 and indices rank is 0..4; rank-0 indices contain one index. Negative axis normalizes by input_rank; negative batch_dims normalizes by coords_rank. After normalization, 0 <= batch_dims <= coords_rank and batch_dims <= axis < input_rank. Leading batch dimensions must match. The inferred output shape is input_shape[:axis] + indices_shape[batch_dims:] + input_shape[axis + 1:].\n\nShapes use the first rank fields of `cmsis_nn_dims` in n, h, w, c order; unused fields are ignored. The inferred output rank must be 0..4, and output_dims must match its leading dimensions. A rank-0 output is one element. All dimension extents must be nonnegative. Input, index and output buffer byte counts must each fit INT32_MAX; this is a CORE capacity limit.\n\nAll metadata pointers are required. A NULL data, indices or output pointer is accepted only when that respective buffer has zero elements. Valid empty calls copy nothing. All supplied indices are checked, even for empty output. Coordinates must be nonnegative and below their corresponding axis extent. Invalid metadata or indices return ARG_ERROR without changing output.\n\nThis operation preserves all bits, including NaN payloads, signed zero and subnormals, independently of floating-point controls. Buffers must not overlap. No scratch buffer is required. Portable copies use existing MVE copy paths when enabled.",
              "examples": [],
              "id": "arm_gather_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gather_f16",
              "params": [
                {
                  "description": "Input data buffer.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Input shape in leading-dimension order.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Signed 32-bit indices.",
                  "direction": "in",
                  "name": "indices_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Indices shape in leading-dimension order.",
                  "direction": "in",
                  "name": "indices_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Ranks and gathering parameters.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gather_params *"
                },
                {
                  "description": "Output data buffer.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Inferred output shape in leading-dimension order.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gather_f16(\n    const float16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t *indices_data,\n    const cmsis_nn_dims *indices_dims,\n    const cmsis_nn_gather_params *params,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 3866,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3866"
              },
              "summary": "Gather contiguous slices along an axis."
            },
            {
              "description": "Gather contiguous slices using coordinate tuples.\n\nData rank is 1..4 and indices rank is 1..4. The final indices dimension is the tuple width, which must be at least one. batch_dims is a TensorFlow-style extension (not a LiteRT builtin option): 0 <= batch_dims < indices_rank, batch_dims < params_rank, and batch_dims + tuple_width <= params_rank. Leading batch dimensions must match. The inferred output shape is indices_shape[:-1] + params_shape[batch_dims + tuple_width:]. Empty data with a nonempty index buffer is rejected.\n\nShapes use the first rank fields of `cmsis_nn_dims` in n, h, w, c order; unused fields are ignored. The inferred output rank must be 0..4, and output_dims must match its leading dimensions. A rank-0 output is one element. All dimension extents must be nonnegative. Input, index and output buffer byte counts must each fit INT32_MAX; this is a CORE capacity limit.\n\nAll metadata pointers are required. A NULL data, indices or output pointer is accepted only when that respective buffer has zero elements. Valid empty calls copy nothing. All supplied indices are checked, even for empty output. Coordinates must be nonnegative and below their corresponding axis extent. Invalid metadata or indices return ARG_ERROR without changing output.\n\nThis operation preserves all bits, including NaN payloads, signed zero and subnormals, independently of floating-point controls. Buffers must not overlap. No scratch buffer is required. Portable copies use existing MVE copy paths when enabled.",
              "examples": [],
              "id": "arm_gather_nd_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gather_nd_f16",
              "params": [
                {
                  "description": "Input data buffer.",
                  "direction": "in",
                  "name": "params_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Input shape in leading-dimension order.",
                  "direction": "in",
                  "name": "params_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Signed 32-bit indices.",
                  "direction": "in",
                  "name": "indices_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Indices shape in leading-dimension order.",
                  "direction": "in",
                  "name": "indices_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Ranks and gathering parameters.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gather_nd_params *"
                },
                {
                  "description": "Output data buffer.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Inferred output shape in leading-dimension order.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gather_nd_f16(\n    const float16_t *params_data,\n    const cmsis_nn_dims *params_dims,\n    const int32_t *indices_data,\n    const cmsis_nn_dims *indices_dims,\n    const cmsis_nn_gather_nd_params *params,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 3911,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3911"
              },
              "summary": "Gather contiguous slices using coordinate tuples."
            }
          ]
        },
        {
          "description": "Data structure types used by private functions.",
          "name": "Structure Types",
          "path": "heliaCORE.genPrivTypes",
          "submodules": [],
          "summary": "Data structure types used by private functions.",
          "symbols": [
            {
              "description": "Union for SIMD access of q31/s16/s8 types.",
              "examples": [],
              "id": "arm_nnword",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "q31 type",
                  "examples": [],
                  "id": "arm_nnword::word",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "word",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t word",
                  "source": {
                    "line": 598,
                    "path": "Include/arm_nnsupportfunctions.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L598"
                  },
                  "summary": "q31 type"
                },
                {
                  "description": "s16 type",
                  "examples": [],
                  "id": "arm_nnword::half_words",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "half_words",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int16_t half_words[2]",
                  "source": {
                    "line": 600,
                    "path": "Include/arm_nnsupportfunctions.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L600"
                  },
                  "summary": "s16 type"
                },
                {
                  "description": "s8 type",
                  "examples": [],
                  "id": "arm_nnword::bytes",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "bytes",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int8_t bytes[4]",
                  "source": {
                    "line": 602,
                    "path": "Include/arm_nnsupportfunctions.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L602"
                  },
                  "summary": "s8 type"
                }
              ],
              "name": "arm_nnword",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "union arm_nnword",
              "source": {
                "line": 596,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L596"
              },
              "summary": "Union for SIMD access of q31/s16/s8 types."
            },
            {
              "description": "Union for data type long long.",
              "examples": [],
              "id": "arm_nn_double",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_double::low",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "low",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "uint32_t low",
                  "source": {
                    "line": 611,
                    "path": "Include/arm_nnsupportfunctions.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L611"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_double::high",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "high",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t high",
                  "source": {
                    "line": 612,
                    "path": "Include/arm_nnsupportfunctions.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L612"
                  },
                  "summary": ""
                }
              ],
              "name": "arm_nn_double",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct arm_nn_double",
              "source": {
                "line": 609,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L609"
              },
              "summary": "Union for data type long long."
            },
            {
              "description": "",
              "examples": [],
              "id": "arm_nn_long_long",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_long_long::long_long",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "long_long",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int64_t long_long",
                  "source": {
                    "line": 617,
                    "path": "Include/arm_nnsupportfunctions.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L617"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_long_long::word",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "word",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "struct arm_nn_double word",
                  "source": {
                    "line": 618,
                    "path": "Include/arm_nnsupportfunctions.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L618"
                  },
                  "summary": ""
                }
              ],
              "name": "arm_nn_long_long",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "union arm_nn_long_long",
              "source": {
                "line": 615,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L615"
              },
              "summary": ""
            }
          ]
        },
        {
          "description": "Enums and Data Structures used in public API.",
          "name": "Structure Types",
          "path": "heliaCORE.genPubTypes",
          "submodules": [],
          "summary": "Enums and Data Structures used in public API.",
          "symbols": [
            {
              "description": "CMSIS-NN object to contain the width and height of a tile",
              "examples": [],
              "id": "cmsis_nn_tile",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Width",
                  "examples": [],
                  "id": "cmsis_nn_tile::w",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "w",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t w",
                  "source": {
                    "line": 107,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L107"
                  },
                  "summary": "Width"
                },
                {
                  "description": "Height",
                  "examples": [],
                  "id": "cmsis_nn_tile::h",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "h",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t h",
                  "source": {
                    "line": 108,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L108"
                  },
                  "summary": "Height"
                }
              ],
              "name": "cmsis_nn_tile",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_tile",
              "source": {
                "line": 105,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L105"
              },
              "summary": "CMSIS-NN object to contain the width and height of a tile"
            },
            {
              "description": "CMSIS-NN object used for the function context.",
              "examples": [],
              "id": "cmsis_nn_context",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Pointer to a buffer needed for the optimization",
                  "examples": [],
                  "id": "cmsis_nn_context::buf",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "buf",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "void * buf",
                  "source": {
                    "line": 114,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L114"
                  },
                  "summary": "Pointer to a buffer needed for the optimization"
                },
                {
                  "description": "Buffer size",
                  "examples": [],
                  "id": "cmsis_nn_context::size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t size",
                  "source": {
                    "line": 115,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L115"
                  },
                  "summary": "Buffer size"
                }
              ],
              "name": "cmsis_nn_context",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_context",
              "source": {
                "line": 112,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L112"
              },
              "summary": "CMSIS-NN object used for the function context."
            },
            {
              "description": "CMSIS-NN object used to hold bias data for int16 variants.",
              "examples": [],
              "id": "cmsis_nn_bias_data",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Pointer to bias data",
                  "examples": [],
                  "id": "cmsis_nn_bias_data::data",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "data",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const void * data",
                  "source": {
                    "line": 121,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L121"
                  },
                  "summary": "Pointer to bias data"
                },
                {
                  "description": "Indicate type of bias data. True means int32 else int64",
                  "examples": [],
                  "id": "cmsis_nn_bias_data::is_int32_bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "is_int32_bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const bool is_int32_bias",
                  "source": {
                    "line": 122,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L122"
                  },
                  "summary": "Indicate type of bias data."
                }
              ],
              "name": "cmsis_nn_bias_data",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_bias_data",
              "source": {
                "line": 119,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L119"
              },
              "summary": "CMSIS-NN object used to hold bias data for int16 variants."
            },
            {
              "description": "CMSIS-NN object to contain the dimensions of the tensors",
              "examples": [],
              "id": "cmsis_nn_dims",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Generic dimension to contain either the batch size or output channels. Please refer to the function documentation for more information",
                  "examples": [],
                  "id": "cmsis_nn_dims::n",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "n",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t n",
                  "source": {
                    "line": 128,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L128"
                  },
                  "summary": "Generic dimension to contain either the batch size or output channels."
                },
                {
                  "description": "Height",
                  "examples": [],
                  "id": "cmsis_nn_dims::h",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "h",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t h",
                  "source": {
                    "line": 130,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L130"
                  },
                  "summary": "Height"
                },
                {
                  "description": "Width",
                  "examples": [],
                  "id": "cmsis_nn_dims::w",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "w",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t w",
                  "source": {
                    "line": 131,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L131"
                  },
                  "summary": "Width"
                },
                {
                  "description": "Input channels",
                  "examples": [],
                  "id": "cmsis_nn_dims::c",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "c",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t c",
                  "source": {
                    "line": 132,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L132"
                  },
                  "summary": "Input channels"
                }
              ],
              "name": "cmsis_nn_dims",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_dims",
              "source": {
                "line": 126,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L126"
              },
              "summary": "CMSIS-NN object to contain the dimensions of the tensors"
            },
            {
              "description": "CMSIS-NN object to contain LSTM specific input parameters related to dimensions",
              "examples": [],
              "id": "cmsis_nn_lstm_dims",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_dims::max_time",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "max_time",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t max_time",
                  "source": {
                    "line": 138,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L138"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_dims::num_inputs",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "num_inputs",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t num_inputs",
                  "source": {
                    "line": 139,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L139"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_dims::num_batches",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "num_batches",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t num_batches",
                  "source": {
                    "line": 140,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L140"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_dims::num_outputs",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "num_outputs",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t num_outputs",
                  "source": {
                    "line": 141,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L141"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_lstm_dims",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_dims",
              "source": {
                "line": 136,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L136"
              },
              "summary": "CMSIS-NN object to contain LSTM specific input parameters related to dimensions"
            },
            {
              "description": "CMSIS-NN object for the per-channel quantization parameters",
              "examples": [],
              "id": "cmsis_nn_per_channel_quant_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Multiplier values",
                  "examples": [],
                  "id": "cmsis_nn_per_channel_quant_params::multiplier",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "multiplier",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t * multiplier",
                  "source": {
                    "line": 147,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L147"
                  },
                  "summary": "Multiplier values"
                },
                {
                  "description": "Shift values",
                  "examples": [],
                  "id": "cmsis_nn_per_channel_quant_params::shift",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "shift",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t * shift",
                  "source": {
                    "line": 148,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L148"
                  },
                  "summary": "Shift values"
                }
              ],
              "name": "cmsis_nn_per_channel_quant_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_per_channel_quant_params",
              "source": {
                "line": 145,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L145"
              },
              "summary": "CMSIS-NN object for the per-channel quantization parameters"
            },
            {
              "description": "CMSIS-NN object for the per-tensor quantization parameters",
              "examples": [],
              "id": "cmsis_nn_per_tensor_quant_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Multiplier value",
                  "examples": [],
                  "id": "cmsis_nn_per_tensor_quant_params::multiplier",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "multiplier",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t multiplier",
                  "source": {
                    "line": 154,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L154"
                  },
                  "summary": "Multiplier value"
                },
                {
                  "description": "Shift value",
                  "examples": [],
                  "id": "cmsis_nn_per_tensor_quant_params::shift",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "shift",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t shift",
                  "source": {
                    "line": 155,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L155"
                  },
                  "summary": "Shift value"
                }
              ],
              "name": "cmsis_nn_per_tensor_quant_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_per_tensor_quant_params",
              "source": {
                "line": 152,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L152"
              },
              "summary": "CMSIS-NN object for the per-tensor quantization parameters"
            },
            {
              "description": "CMSIS-NN object for quantization parameters. This struct supports both per-tensor and per-channels requantization and is recommended for new operators.",
              "examples": [],
              "id": "cmsis_nn_quant_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Multiplier values",
                  "examples": [],
                  "id": "cmsis_nn_quant_params::multiplier",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "multiplier",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t * multiplier",
                  "source": {
                    "line": 164,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L164"
                  },
                  "summary": "Multiplier values"
                },
                {
                  "description": "Shift values",
                  "examples": [],
                  "id": "cmsis_nn_quant_params::shift",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "shift",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t * shift",
                  "source": {
                    "line": 165,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L165"
                  },
                  "summary": "Shift values"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_quant_params::is_per_channel",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "is_per_channel",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t is_per_channel",
                  "source": {
                    "line": 166,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L166"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_quant_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_quant_params",
              "source": {
                "line": 162,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L162"
              },
              "summary": "CMSIS-NN object for quantization parameters."
            },
            {
              "description": "CMSIS-NN object for the quantized Relu activation",
              "examples": [],
              "id": "cmsis_nn_activation",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Min value used to clamp the result",
                  "examples": [],
                  "id": "cmsis_nn_activation::min",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "min",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t min",
                  "source": {
                    "line": 172,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L172"
                  },
                  "summary": "Min value used to clamp the result"
                },
                {
                  "description": "Max value used to clamp the result",
                  "examples": [],
                  "id": "cmsis_nn_activation::max",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "max",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t max",
                  "source": {
                    "line": 173,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L173"
                  },
                  "summary": "Max value used to clamp the result"
                }
              ],
              "name": "cmsis_nn_activation",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_activation",
              "source": {
                "line": 170,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L170"
              },
              "summary": "CMSIS-NN object for the quantized Relu activation"
            },
            {
              "description": "CMSIS-NN object for the convolution layer parameters",
              "examples": [],
              "id": "cmsis_nn_conv_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "The negative of the zero value for the input tensor",
                  "examples": [],
                  "id": "cmsis_nn_conv_params::input_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_offset",
                  "source": {
                    "line": 179,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L179"
                  },
                  "summary": "The negative of the zero value for the input tensor"
                },
                {
                  "description": "The negative of the zero value for the output tensor",
                  "examples": [],
                  "id": "cmsis_nn_conv_params::output_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t output_offset",
                  "source": {
                    "line": 180,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L180"
                  },
                  "summary": "The negative of the zero value for the output tensor"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_conv_params::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 181,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L181"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_conv_params::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 182,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L182"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_conv_params::dilation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "dilation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile dilation",
                  "source": {
                    "line": 183,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L183"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_conv_params::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation activation",
                  "source": {
                    "line": 184,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L184"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_conv_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_conv_params",
              "source": {
                "line": 177,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L177"
              },
              "summary": "CMSIS-NN object for the convolution layer parameters"
            },
            {
              "description": "CMSIS-NN object for the transpose convolution layer parameters",
              "examples": [],
              "id": "cmsis_nn_transpose_conv_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "The negative of the zero value for the input tensor",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params::input_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_offset",
                  "source": {
                    "line": 190,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L190"
                  },
                  "summary": "The negative of the zero value for the input tensor"
                },
                {
                  "description": "The negative of the zero value for the output tensor",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params::output_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t output_offset",
                  "source": {
                    "line": 191,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L191"
                  },
                  "summary": "The negative of the zero value for the output tensor"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 192,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L192"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 193,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L193"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params::padding_offsets",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding_offsets",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding_offsets",
                  "source": {
                    "line": 194,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L194"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params::dilation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "dilation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile dilation",
                  "source": {
                    "line": 195,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L195"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation activation",
                  "source": {
                    "line": 196,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L196"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_transpose_conv_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_transpose_conv_params",
              "source": {
                "line": 188,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L188"
              },
              "summary": "CMSIS-NN object for the transpose convolution layer parameters"
            },
            {
              "description": "CMSIS-NN object for the depthwise convolution layer parameters",
              "examples": [],
              "id": "cmsis_nn_dw_conv_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "The negative of the zero value for the input tensor",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params::input_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_offset",
                  "source": {
                    "line": 202,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L202"
                  },
                  "summary": "The negative of the zero value for the input tensor"
                },
                {
                  "description": "The negative of the zero value for the output tensor",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params::output_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t output_offset",
                  "source": {
                    "line": 203,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L203"
                  },
                  "summary": "The negative of the zero value for the output tensor"
                },
                {
                  "description": "Channel Multiplier. ch_mult * in_ch = out_ch",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params::ch_mult",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "ch_mult",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t ch_mult",
                  "source": {
                    "line": 204,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L204"
                  },
                  "summary": "Channel Multiplier."
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 205,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L205"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 206,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L206"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params::dilation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "dilation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile dilation",
                  "source": {
                    "line": 207,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L207"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation activation",
                  "source": {
                    "line": 208,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L208"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_dw_conv_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_dw_conv_params",
              "source": {
                "line": 200,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L200"
              },
              "summary": "CMSIS-NN object for the depthwise convolution layer parameters"
            },
            {
              "description": "CMSIS-NN object for pooling layer parameters",
              "examples": [],
              "id": "cmsis_nn_pool_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_pool_params::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 214,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L214"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_pool_params::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 215,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L215"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_pool_params::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation activation",
                  "source": {
                    "line": 216,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L216"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_pool_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_pool_params",
              "source": {
                "line": 212,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L212"
              },
              "summary": "CMSIS-NN object for pooling layer parameters"
            },
            {
              "description": "CMSIS-NN object for the gather operator",
              "examples": [],
              "id": "cmsis_nn_gather_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Axis to gather from. Supports negative indexing.",
                  "examples": [],
                  "id": "cmsis_nn_gather_params::axis",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "axis",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t axis",
                  "source": {
                    "line": 222,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L222"
                  },
                  "summary": "Axis to gather from."
                },
                {
                  "description": "Number of leading batch dimensions",
                  "examples": [],
                  "id": "cmsis_nn_gather_params::batch_dims",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "batch_dims",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t batch_dims",
                  "source": {
                    "line": 223,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L223"
                  },
                  "summary": "Number of leading batch dimensions"
                },
                {
                  "description": "Rank of the input tensor (range: [1, 4])",
                  "examples": [],
                  "id": "cmsis_nn_gather_params::input_rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_rank",
                  "source": {
                    "line": 224,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L224"
                  },
                  "summary": "Rank of the input tensor (range: [1, 4])"
                },
                {
                  "description": "Rank of the coordinate tensor ([1, 4]; float gather also accepts scalar rank 0)",
                  "examples": [],
                  "id": "cmsis_nn_gather_params::coords_rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "coords_rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t coords_rank",
                  "source": {
                    "line": 225,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L225"
                  },
                  "summary": "Rank of the coordinate tensor ([1, 4]; float gather also accepts scalar rank 0)"
                }
              ],
              "name": "cmsis_nn_gather_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_gather_params",
              "source": {
                "line": 220,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L220"
              },
              "summary": "CMSIS-NN object for the gather operator"
            },
            {
              "description": "CMSIS-NN object for the gather_nd operator",
              "examples": [],
              "id": "cmsis_nn_gather_nd_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Rank of the params tensor (range: [1, 4])",
                  "examples": [],
                  "id": "cmsis_nn_gather_nd_params::params_rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "params_rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t params_rank",
                  "source": {
                    "line": 231,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L231"
                  },
                  "summary": "Rank of the params tensor (range: [1, 4])"
                },
                {
                  "description": "Rank of the indices tensor (range: [1, 4])",
                  "examples": [],
                  "id": "cmsis_nn_gather_nd_params::indices_rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "indices_rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t indices_rank",
                  "source": {
                    "line": 232,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L232"
                  },
                  "summary": "Rank of the indices tensor (range: [1, 4])"
                },
                {
                  "description": "Number of batch dimensions",
                  "examples": [],
                  "id": "cmsis_nn_gather_nd_params::batch_dims",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "batch_dims",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t batch_dims",
                  "source": {
                    "line": 233,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L233"
                  },
                  "summary": "Number of batch dimensions"
                }
              ],
              "name": "cmsis_nn_gather_nd_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_gather_nd_params",
              "source": {
                "line": 229,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L229"
              },
              "summary": "CMSIS-NN object for the gathernd operator"
            },
            {
              "description": "CMSIS-NN object for the tile operator",
              "examples": [],
              "id": "cmsis_nn_tile_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Rank of the input tensor (range: [1, 8])",
                  "examples": [],
                  "id": "cmsis_nn_tile_params::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 239,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L239"
                  },
                  "summary": "Rank of the input tensor (range: [1, 8])"
                },
                {
                  "description": "Input shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_tile_params::input_shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * input_shape",
                  "source": {
                    "line": 240,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L240"
                  },
                  "summary": "Input shape array (length = rank)"
                },
                {
                  "description": "Multiples array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_tile_params::multiples",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "multiples",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * multiples",
                  "source": {
                    "line": 241,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L241"
                  },
                  "summary": "Multiples array (length = rank)"
                }
              ],
              "name": "cmsis_nn_tile_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_tile_params",
              "source": {
                "line": 237,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L237"
              },
              "summary": "CMSIS-NN object for the tile operator"
            },
            {
              "description": "CMSIS-NN object for the broadcast_to operator",
              "examples": [],
              "id": "cmsis_nn_broadcast_to_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Rank of input/output tensors (range: [1, 8])",
                  "examples": [],
                  "id": "cmsis_nn_broadcast_to_params::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 247,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L247"
                  },
                  "summary": "Rank of input/output tensors (range: [1, 8])"
                },
                {
                  "description": "Input shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_broadcast_to_params::input_shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * input_shape",
                  "source": {
                    "line": 248,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L248"
                  },
                  "summary": "Input shape array (length = rank)"
                },
                {
                  "description": "Output (broadcast target) shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_broadcast_to_params::output_shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * output_shape",
                  "source": {
                    "line": 249,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L249"
                  },
                  "summary": "Output (broadcast target) shape array (length = rank)"
                }
              ],
              "name": "cmsis_nn_broadcast_to_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_broadcast_to_params",
              "source": {
                "line": 245,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L245"
              },
              "summary": "CMSIS-NN object for the broadcastto operator"
            },
            {
              "description": "CMSIS-NN object for the scatter_nd operator",
              "examples": [],
              "id": "cmsis_nn_scatter_nd_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Number of update slices",
                  "examples": [],
                  "id": "cmsis_nn_scatter_nd_params::num_updates",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "num_updates",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t num_updates",
                  "source": {
                    "line": 255,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L255"
                  },
                  "summary": "Number of update slices"
                },
                {
                  "description": "Depth of each index vector",
                  "examples": [],
                  "id": "cmsis_nn_scatter_nd_params::index_depth",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "index_depth",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t index_depth",
                  "source": {
                    "line": 256,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L256"
                  },
                  "summary": "Depth of each index vector"
                },
                {
                  "description": "Size of each update slice",
                  "examples": [],
                  "id": "cmsis_nn_scatter_nd_params::slice_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "slice_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t slice_size",
                  "source": {
                    "line": 257,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L257"
                  },
                  "summary": "Size of each update slice"
                },
                {
                  "description": "Total number of elements in output",
                  "examples": [],
                  "id": "cmsis_nn_scatter_nd_params::output_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t output_size",
                  "source": {
                    "line": 258,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L258"
                  },
                  "summary": "Total number of elements in output"
                },
                {
                  "description": "Strides of the output tensor (length = index_depth)",
                  "examples": [],
                  "id": "cmsis_nn_scatter_nd_params::output_strides",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_strides",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * output_strides",
                  "source": {
                    "line": 259,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L259"
                  },
                  "summary": "Strides of the output tensor (length = indexdepth)"
                }
              ],
              "name": "cmsis_nn_scatter_nd_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_scatter_nd_params",
              "source": {
                "line": 253,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L253"
              },
              "summary": "CMSIS-NN object for the scatternd operator"
            },
            {
              "description": "CMSIS-NN object for the mirror_pad operator",
              "examples": [],
              "id": "cmsis_nn_mirror_pad_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Rank of the input tensor (range: [1, 8])",
                  "examples": [],
                  "id": "cmsis_nn_mirror_pad_params::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 265,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L265"
                  },
                  "summary": "Rank of the input tensor (range: [1, 8])"
                },
                {
                  "description": "Input shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_mirror_pad_params::input_shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * input_shape",
                  "source": {
                    "line": 266,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L266"
                  },
                  "summary": "Input shape array (length = rank)"
                },
                {
                  "description": "Output shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_mirror_pad_params::output_shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * output_shape",
                  "source": {
                    "line": 267,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L267"
                  },
                  "summary": "Output shape array (length = rank)"
                },
                {
                  "description": "Padding before each dimension (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_mirror_pad_params::pad_before",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "pad_before",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * pad_before",
                  "source": {
                    "line": 268,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L268"
                  },
                  "summary": "Padding before each dimension (length = rank)"
                },
                {
                  "description": "0 = REFLECT, 1 = SYMMETRIC",
                  "examples": [],
                  "id": "cmsis_nn_mirror_pad_params::mode",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "mode",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t mode",
                  "source": {
                    "line": 269,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L269"
                  },
                  "summary": "0 = REFLECT, 1 = SYMMETRIC"
                }
              ],
              "name": "cmsis_nn_mirror_pad_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_mirror_pad_params",
              "source": {
                "line": 263,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L263"
              },
              "summary": "CMSIS-NN object for the mirrorpad operator"
            },
            {
              "description": "CMSIS-NN object for the WHERE operator",
              "examples": [],
              "id": "cmsis_nn_where_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Rank of the condition tensor (range: [1, 8])",
                  "examples": [],
                  "id": "cmsis_nn_where_params::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 275,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L275"
                  },
                  "summary": "Rank of the condition tensor (range: [1, 8])"
                },
                {
                  "description": "Condition tensor shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_where_params::shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * shape",
                  "source": {
                    "line": 276,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L276"
                  },
                  "summary": "Condition tensor shape array (length = rank)"
                }
              ],
              "name": "cmsis_nn_where_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_where_params",
              "source": {
                "line": 273,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L273"
              },
              "summary": "CMSIS-NN object for the WHERE operator"
            },
            {
              "description": "CMSIS-NN object for the select_v2 operator (with broadcast)",
              "examples": [],
              "id": "cmsis_nn_select_v2_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Rank of the output tensor (range: [1, 8])",
                  "examples": [],
                  "id": "cmsis_nn_select_v2_params::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 282,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L282"
                  },
                  "summary": "Rank of the output tensor (range: [1, 8])"
                },
                {
                  "description": "Output shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_select_v2_params::output_shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * output_shape",
                  "source": {
                    "line": 283,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L283"
                  },
                  "summary": "Output shape array (length = rank)"
                },
                {
                  "description": "Condition tensor broadcast strides (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_select_v2_params::cond_strides",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cond_strides",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * cond_strides",
                  "source": {
                    "line": 284,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L284"
                  },
                  "summary": "Condition tensor broadcast strides (length = rank)"
                },
                {
                  "description": "X tensor broadcast strides (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_select_v2_params::x_strides",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "x_strides",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * x_strides",
                  "source": {
                    "line": 285,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L285"
                  },
                  "summary": "X tensor broadcast strides (length = rank)"
                },
                {
                  "description": "Y tensor broadcast strides (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_select_v2_params::y_strides",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "y_strides",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * y_strides",
                  "source": {
                    "line": 286,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L286"
                  },
                  "summary": "Y tensor broadcast strides (length = rank)"
                }
              ],
              "name": "cmsis_nn_select_v2_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_select_v2_params",
              "source": {
                "line": 280,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L280"
              },
              "summary": "CMSIS-NN object for the selectv2 operator (with broadcast)"
            },
            {
              "description": "CMSIS-NN object for the reverse_sequence operator",
              "examples": [],
              "id": "cmsis_nn_reverse_sequence_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Rank of the input tensor (range: [1, 8])",
                  "examples": [],
                  "id": "cmsis_nn_reverse_sequence_params::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 292,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L292"
                  },
                  "summary": "Rank of the input tensor (range: [1, 8])"
                },
                {
                  "description": "Input shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_reverse_sequence_params::shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * shape",
                  "source": {
                    "line": 293,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L293"
                  },
                  "summary": "Input shape array (length = rank)"
                },
                {
                  "description": "Dimension along which to reverse",
                  "examples": [],
                  "id": "cmsis_nn_reverse_sequence_params::seq_dim",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "seq_dim",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t seq_dim",
                  "source": {
                    "line": 294,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L294"
                  },
                  "summary": "Dimension along which to reverse"
                },
                {
                  "description": "Batch dimension",
                  "examples": [],
                  "id": "cmsis_nn_reverse_sequence_params::batch_dim",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "batch_dim",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t batch_dim",
                  "source": {
                    "line": 295,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L295"
                  },
                  "summary": "Batch dimension"
                }
              ],
              "name": "cmsis_nn_reverse_sequence_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_reverse_sequence_params",
              "source": {
                "line": 290,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L290"
              },
              "summary": "CMSIS-NN object for the reversesequence operator"
            },
            {
              "description": "CMSIS-NN object for the dynamic_update_slice operator",
              "examples": [],
              "id": "cmsis_nn_dynamic_update_slice_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Rank of the operand tensor (range: [1, 8])",
                  "examples": [],
                  "id": "cmsis_nn_dynamic_update_slice_params::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 301,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L301"
                  },
                  "summary": "Rank of the operand tensor (range: [1, 8])"
                },
                {
                  "description": "Operand shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_dynamic_update_slice_params::operand_shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "operand_shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * operand_shape",
                  "source": {
                    "line": 302,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L302"
                  },
                  "summary": "Operand shape array (length = rank)"
                },
                {
                  "description": "Update shape array (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_dynamic_update_slice_params::update_shape",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "update_shape",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * update_shape",
                  "source": {
                    "line": 303,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L303"
                  },
                  "summary": "Update shape array (length = rank)"
                },
                {
                  "description": "Total number of elements in operand",
                  "examples": [],
                  "id": "cmsis_nn_dynamic_update_slice_params::operand_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "operand_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t operand_size",
                  "source": {
                    "line": 304,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L304"
                  },
                  "summary": "Total number of elements in operand"
                },
                {
                  "description": "Total number of elements in update",
                  "examples": [],
                  "id": "cmsis_nn_dynamic_update_slice_params::update_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "update_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t update_size",
                  "source": {
                    "line": 305,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L305"
                  },
                  "summary": "Total number of elements in update"
                },
                {
                  "description": "Strides of the operand tensor (length = rank)",
                  "examples": [],
                  "id": "cmsis_nn_dynamic_update_slice_params::operand_strides",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "operand_strides",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t * operand_strides",
                  "source": {
                    "line": 306,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L306"
                  },
                  "summary": "Strides of the operand tensor (length = rank)"
                }
              ],
              "name": "cmsis_nn_dynamic_update_slice_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_dynamic_update_slice_params",
              "source": {
                "line": 299,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L299"
              },
              "summary": "CMSIS-NN object for the dynamicupdateslice operator"
            },
            {
              "description": "CMSIS-NN object for Fully Connected layer parameters",
              "examples": [],
              "id": "cmsis_nn_fc_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "The negative of the zero value for the input tensor",
                  "examples": [],
                  "id": "cmsis_nn_fc_params::input_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_offset",
                  "source": {
                    "line": 312,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L312"
                  },
                  "summary": "The negative of the zero value for the input tensor"
                },
                {
                  "description": "The negative of the zero value for the filter tensor",
                  "examples": [],
                  "id": "cmsis_nn_fc_params::filter_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "filter_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t filter_offset",
                  "source": {
                    "line": 313,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L313"
                  },
                  "summary": "The negative of the zero value for the filter tensor"
                },
                {
                  "description": "The negative of the zero value for the output tensor",
                  "examples": [],
                  "id": "cmsis_nn_fc_params::output_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t output_offset",
                  "source": {
                    "line": 314,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L314"
                  },
                  "summary": "The negative of the zero value for the output tensor"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_fc_params::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation activation",
                  "source": {
                    "line": 315,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L315"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_fc_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_fc_params",
              "source": {
                "line": 310,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L310"
              },
              "summary": "CMSIS-NN object for Fully Connected layer parameters"
            },
            {
              "description": "CMSIS-NN object for Batch Matmul layer parameters",
              "examples": [],
              "id": "cmsis_nn_bmm_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params::adj_x",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "adj_x",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const bool adj_x",
                  "source": {
                    "line": 321,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L321"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params::adj_y",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "adj_y",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const bool adj_y",
                  "source": {
                    "line": 322,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L322"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params::fc_params",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "fc_params",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_fc_params fc_params",
                  "source": {
                    "line": 323,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L323"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_bmm_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_bmm_params",
              "source": {
                "line": 319,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L319"
              },
              "summary": "CMSIS-NN object for Batch Matmul layer parameters"
            },
            {
              "description": "CMSIS-NN object for Transpose layer parameters",
              "examples": [],
              "id": "cmsis_nn_transpose_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_transpose_params::num_dims",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "num_dims",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t num_dims",
                  "source": {
                    "line": 329,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L329"
                  },
                  "summary": ""
                },
                {
                  "description": "The dimensions applied to the input dimensions",
                  "examples": [],
                  "id": "cmsis_nn_transpose_params::permutations",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "permutations",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const uint32_t * permutations",
                  "source": {
                    "line": 330,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L330"
                  },
                  "summary": "The dimensions applied to the input dimensions"
                }
              ],
              "name": "cmsis_nn_transpose_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_transpose_params",
              "source": {
                "line": 327,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L327"
              },
              "summary": "CMSIS-NN object for Transpose layer parameters"
            },
            {
              "description": "CMSIS-NN object for Resize Nearest Neighbor layer parameters",
              "examples": [],
              "id": "cmsis_nn_resize_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Align corners when calculating interpolation",
                  "examples": [],
                  "id": "cmsis_nn_resize_params::align_corners",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "align_corners",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "bool align_corners",
                  "source": {
                    "line": 336,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L336"
                  },
                  "summary": "Align corners when calculating interpolation"
                },
                {
                  "description": "Use half pixel centers when calculating interpolation",
                  "examples": [],
                  "id": "cmsis_nn_resize_params::half_pixel_centers",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "half_pixel_centers",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "bool half_pixel_centers",
                  "source": {
                    "line": 337,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L337"
                  },
                  "summary": "Use half pixel centers when calculating interpolation"
                }
              ],
              "name": "cmsis_nn_resize_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_resize_params",
              "source": {
                "line": 334,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L334"
              },
              "summary": "CMSIS-NN object for Resize Nearest Neighbor layer parameters"
            },
            {
              "description": "CMSIS-NN object for SVDF layer parameters",
              "examples": [],
              "id": "cmsis_nn_svdf_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 343,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L343"
                  },
                  "summary": ""
                },
                {
                  "description": "The negative of the zero value for the input tensor",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params::input_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_offset",
                  "source": {
                    "line": 344,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L344"
                  },
                  "summary": "The negative of the zero value for the input tensor"
                },
                {
                  "description": "The negative of the zero value for the output tensor",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params::output_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t output_offset",
                  "source": {
                    "line": 345,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L345"
                  },
                  "summary": "The negative of the zero value for the output tensor"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params::input_activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation input_activation",
                  "source": {
                    "line": 346,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L346"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params::output_activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation output_activation",
                  "source": {
                    "line": 347,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L347"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_svdf_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_svdf_params",
              "source": {
                "line": 341,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L341"
              },
              "summary": "CMSIS-NN object for SVDF layer parameters"
            },
            {
              "description": "CMSIS-NN object for Softmax s16 layer parameters",
              "examples": [],
              "id": "cmsis_nn_softmax_lut_s16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_softmax_lut_s16::exp_lut",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "exp_lut",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int16_t * exp_lut",
                  "source": {
                    "line": 353,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L353"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_softmax_lut_s16::one_by_one_lut",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "one_by_one_lut",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int16_t * one_by_one_lut",
                  "source": {
                    "line": 354,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L354"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_softmax_lut_s16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_softmax_lut_s16",
              "source": {
                "line": 351,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L351"
              },
              "summary": "CMSIS-NN object for Softmax s16 layer parameters"
            },
            {
              "description": "",
              "examples": [],
              "id": "cmsis_nn_concatenation_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_concatenation_params::axis",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "axis",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const int32_t axis",
                  "source": {
                    "line": 359,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L359"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_concatenation_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_concatenation_params",
              "source": {
                "line": 357,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L357"
              },
              "summary": ""
            },
            {
              "description": "CMSIS-NN object for quantization parameters",
              "examples": [],
              "id": "cmsis_nn_scaling",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Multiplier value",
                  "examples": [],
                  "id": "cmsis_nn_scaling::multiplier",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "multiplier",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t multiplier",
                  "source": {
                    "line": 365,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L365"
                  },
                  "summary": "Multiplier value"
                },
                {
                  "description": "Shift value",
                  "examples": [],
                  "id": "cmsis_nn_scaling::shift",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "shift",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t shift",
                  "source": {
                    "line": 366,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L366"
                  },
                  "summary": "Shift value"
                }
              ],
              "name": "cmsis_nn_scaling",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_scaling",
              "source": {
                "line": 363,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L363"
              },
              "summary": "CMSIS-NN object for quantization parameters"
            },
            {
              "description": "CMSIS-NN object for LSTM gate parameters",
              "examples": [],
              "id": "cmsis_nn_lstm_gate",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::input_multiplier",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_multiplier",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_multiplier",
                  "source": {
                    "line": 372,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L372"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::input_shift",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_shift",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_shift",
                  "source": {
                    "line": 373,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L373"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::input_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const void * input_weights",
                  "source": {
                    "line": 374,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L374"
                  },
                  "summary": ""
                },
                {
                  "description": "Bias added with precomputed kernel_sum * lhs_offset",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::input_effective_bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_effective_bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const void * input_effective_bias",
                  "source": {
                    "line": 375,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L375"
                  },
                  "summary": "Bias added with precomputed kernelsum  lhsoffset"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::hidden_multiplier",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_multiplier",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t hidden_multiplier",
                  "source": {
                    "line": 377,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L377"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::hidden_shift",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_shift",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t hidden_shift",
                  "source": {
                    "line": 378,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L378"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::hidden_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const void * hidden_weights",
                  "source": {
                    "line": 379,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L379"
                  },
                  "summary": ""
                },
                {
                  "description": "Precomputed kernel_sum * lhs_offset",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::hidden_effective_bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_effective_bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const void * hidden_effective_bias",
                  "source": {
                    "line": 380,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L380"
                  },
                  "summary": "Precomputed kernelsum  lhsoffset"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const void * bias",
                  "source": {
                    "line": 382,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L382"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate::activation_type",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation_type",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_activation_type activation_type",
                  "source": {
                    "line": 383,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L383"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_lstm_gate",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_gate",
              "source": {
                "line": 370,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L370"
              },
              "summary": "CMSIS-NN object for LSTM gate parameters"
            },
            {
              "description": "CMSIS-NN object for LSTM parameters",
              "examples": [],
              "id": "cmsis_nn_lstm_params",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "0 if first dimension is batch, else first dimension is time",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::time_major",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_major",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_major",
                  "source": {
                    "line": 389,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L389"
                  },
                  "summary": "0 if first dimension is batch, else first dimension is time"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::batch_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "batch_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t batch_size",
                  "source": {
                    "line": 390,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L390"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::time_steps",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_steps",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_steps",
                  "source": {
                    "line": 391,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L391"
                  },
                  "summary": ""
                },
                {
                  "description": "Size of new data input into the LSTM cell",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::input_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_size",
                  "source": {
                    "line": 392,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L392"
                  },
                  "summary": "Size of new data input into the LSTM cell"
                },
                {
                  "description": "Size of output from the LSTM cell, used as output and recursively into the next time step",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::hidden_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t hidden_size",
                  "source": {
                    "line": 394,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L394"
                  },
                  "summary": "Size of output from the LSTM cell, used as output and recursively into the next time step"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::input_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_offset",
                  "source": {
                    "line": 396,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L396"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::forget_to_cell_multiplier",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "forget_to_cell_multiplier",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t forget_to_cell_multiplier",
                  "source": {
                    "line": 398,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L398"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::forget_to_cell_shift",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "forget_to_cell_shift",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t forget_to_cell_shift",
                  "source": {
                    "line": 399,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L399"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::input_to_cell_multiplier",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_to_cell_multiplier",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_to_cell_multiplier",
                  "source": {
                    "line": 400,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L400"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::input_to_cell_shift",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_to_cell_shift",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_to_cell_shift",
                  "source": {
                    "line": 401,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L401"
                  },
                  "summary": ""
                },
                {
                  "description": "Min/max value of cell output",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::cell_clip",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_clip",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t cell_clip",
                  "source": {
                    "line": 402,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L402"
                  },
                  "summary": "Min/max value of cell output"
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::cell_scale_power",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_scale_power",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t cell_scale_power",
                  "source": {
                    "line": 403,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L403"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::output_multiplier",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_multiplier",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t output_multiplier",
                  "source": {
                    "line": 405,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L405"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::output_shift",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_shift",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t output_shift",
                  "source": {
                    "line": 406,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L406"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::output_offset",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_offset",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t output_offset",
                  "source": {
                    "line": 407,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L407"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::forget_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "forget_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate forget_gate",
                  "source": {
                    "line": 409,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L409"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::input_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate input_gate",
                  "source": {
                    "line": 410,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L410"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::cell_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate cell_gate",
                  "source": {
                    "line": 411,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L411"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params::output_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate output_gate",
                  "source": {
                    "line": 412,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L412"
                  },
                  "summary": ""
                }
              ],
              "name": "cmsis_nn_lstm_params",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_params",
              "source": {
                "line": 387,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L387"
              },
              "summary": "CMSIS-NN object for LSTM parameters"
            },
            {
              "description": "CMSIS-NN object for LSTM scratch buffers.\n\nThere is no size field and no runtime enforcement: an undersized temp1 or temp2 is written past on every build target, so size them from the queries below, not by transcribing a formula.",
              "examples": [],
              "id": "cmsis_nn_lstm_context",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Gate-vector scratch (int16_t elements for both the s8 and s16 layers). Sized by `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` / `arm_lstm_unidirectional_s16_temp1_get_buffer_size()`.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context::temp1",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "temp1",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "void * temp1",
                  "source": {
                    "line": 422,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L422"
                  },
                  "summary": "Gate-vector scratch (int16t elements for both the s8 and s16 layers)."
                },
                {
                  "description": "Cell-gate and tanh(cell_state) scratch (int16_t elements for both layers). Sized by `arm_lstm_unidirectional_s8_temp2_get_buffer_size()` / `arm_lstm_unidirectional_s16_temp2_get_buffer_size()`.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context::temp2",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "temp2",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "void * temp2",
                  "source": {
                    "line": 425,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L425"
                  },
                  "summary": "Cell-gate and tanh(cellstate) scratch (int16t elements for both layers)."
                },
                {
                  "description": "Cell-state buffer, batch_size * hidden_size int16_t elements for both layers.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context::cell_state",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_state",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "void * cell_state",
                  "source": {
                    "line": 428,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L428"
                  },
                  "summary": "Cell-state buffer, batchsize  hiddensize int16t elements for both layers."
                },
                {
                  "description": "Optional in/out hidden state for streaming; NULL selects stateless.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context::hidden_state",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_state",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "void * hidden_state",
                  "source": {
                    "line": 429,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L429"
                  },
                  "summary": "Optional in/out hidden state for streaming; NULL selects stateless."
                }
              ],
              "name": "cmsis_nn_lstm_context",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_context",
              "source": {
                "line": 420,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L420"
              },
              "summary": "CMSIS-NN object for LSTM scratch buffers."
            },
            {
              "description": "Activation clamp range for floating-point operators.",
              "examples": [],
              "id": "cmsis_nn_activation_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Minimum value used to clamp the result.",
                  "examples": [],
                  "id": "cmsis_nn_activation_f32::min",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "min",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float32_t min",
                  "source": {
                    "line": 135,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L135"
                  },
                  "summary": "Minimum value used to clamp the result."
                },
                {
                  "description": "Maximum value used to clamp the result.",
                  "examples": [],
                  "id": "cmsis_nn_activation_f32::max",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "max",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float32_t max",
                  "source": {
                    "line": 136,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L136"
                  },
                  "summary": "Maximum value used to clamp the result."
                }
              ],
              "name": "cmsis_nn_activation_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_activation_f32",
              "source": {
                "line": 133,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L133"
              },
              "summary": "Activation clamp range for floating-point operators."
            },
            {
              "description": "Convolution parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_conv_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Spatial stride.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f32::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 144,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L144"
                  },
                  "summary": "Spatial stride."
                },
                {
                  "description": "Spatial zero-padding.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f32::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 145,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L145"
                  },
                  "summary": "Spatial zero-padding."
                },
                {
                  "description": "Spatial dilation.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f32::dilation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "dilation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile dilation",
                  "source": {
                    "line": 146,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L146"
                  },
                  "summary": "Spatial dilation."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f32::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f32 activation",
                  "source": {
                    "line": 147,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L147"
                  },
                  "summary": "Output activation clamp range."
                },
                {
                  "description": "Filter storage format.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f32::weight_format",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "weight_format",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_weight_format_flt weight_format",
                  "source": {
                    "line": 148,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L148"
                  },
                  "summary": "Filter storage format."
                }
              ],
              "name": "cmsis_nn_conv_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_conv_params_f32",
              "source": {
                "line": 142,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L142"
              },
              "summary": "Convolution parameters for float32 operators."
            },
            {
              "description": "Transpose convolution parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_transpose_conv_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Spatial stride.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f32::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 156,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L156"
                  },
                  "summary": "Spatial stride."
                },
                {
                  "description": "Spatial zero-padding.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f32::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 157,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L157"
                  },
                  "summary": "Spatial zero-padding."
                },
                {
                  "description": "Output padding adjustment for transpose convolution.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f32::padding_offsets",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding_offsets",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding_offsets",
                  "source": {
                    "line": 158,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L158"
                  },
                  "summary": "Output padding adjustment for transpose convolution."
                },
                {
                  "description": "Spatial dilation.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f32::dilation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "dilation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile dilation",
                  "source": {
                    "line": 159,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L159"
                  },
                  "summary": "Spatial dilation."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f32::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f32 activation",
                  "source": {
                    "line": 160,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L160"
                  },
                  "summary": "Output activation clamp range."
                }
              ],
              "name": "cmsis_nn_transpose_conv_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_transpose_conv_params_f32",
              "source": {
                "line": 154,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L154"
              },
              "summary": "Transpose convolution parameters for float32 operators."
            },
            {
              "description": "Depthwise convolution parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_dw_conv_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Channel multiplier. `ch_mult * in_ch = out_ch`.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f32::ch_mult",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "ch_mult",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t ch_mult",
                  "source": {
                    "line": 168,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L168"
                  },
                  "summary": "Channel multiplier."
                },
                {
                  "description": "Spatial stride.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f32::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 169,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L169"
                  },
                  "summary": "Spatial stride."
                },
                {
                  "description": "Spatial zero-padding.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f32::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 170,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L170"
                  },
                  "summary": "Spatial zero-padding."
                },
                {
                  "description": "Spatial dilation.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f32::dilation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "dilation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile dilation",
                  "source": {
                    "line": 171,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L171"
                  },
                  "summary": "Spatial dilation."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f32::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f32 activation",
                  "source": {
                    "line": 172,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L172"
                  },
                  "summary": "Output activation clamp range."
                }
              ],
              "name": "cmsis_nn_dw_conv_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_dw_conv_params_f32",
              "source": {
                "line": 166,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L166"
              },
              "summary": "Depthwise convolution parameters for float32 operators."
            },
            {
              "description": "Pooling parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_pool_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Spatial stride.",
                  "examples": [],
                  "id": "cmsis_nn_pool_params_f32::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 180,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L180"
                  },
                  "summary": "Spatial stride."
                },
                {
                  "description": "Spatial zero-padding.",
                  "examples": [],
                  "id": "cmsis_nn_pool_params_f32::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 181,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L181"
                  },
                  "summary": "Spatial zero-padding."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_pool_params_f32::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f32 activation",
                  "source": {
                    "line": 182,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L182"
                  },
                  "summary": "Output activation clamp range."
                }
              ],
              "name": "cmsis_nn_pool_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_pool_params_f32",
              "source": {
                "line": 178,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L178"
              },
              "summary": "Pooling parameters for float32 operators."
            },
            {
              "description": "Fully connected layer parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_fc_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_fc_params_f32::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f32 activation",
                  "source": {
                    "line": 190,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L190"
                  },
                  "summary": "Output activation clamp range."
                },
                {
                  "description": "Weight storage format.",
                  "examples": [],
                  "id": "cmsis_nn_fc_params_f32::weight_format",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "weight_format",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_weight_format_flt weight_format",
                  "source": {
                    "line": 191,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L191"
                  },
                  "summary": "Weight storage format."
                }
              ],
              "name": "cmsis_nn_fc_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_fc_params_f32",
              "source": {
                "line": 188,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L188"
              },
              "summary": "Fully connected layer parameters for float32 operators."
            },
            {
              "description": "Batched matrix multiplication parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_bmm_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "True when the left-hand-side operand is stored transposed.",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params_f32::adj_x",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "adj_x",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const bool adj_x",
                  "source": {
                    "line": 199,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L199"
                  },
                  "summary": "True when the left-hand-side operand is stored transposed."
                },
                {
                  "description": "True when the right-hand-side operand is stored transposed.",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params_f32::adj_y",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "adj_y",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const bool adj_y",
                  "source": {
                    "line": 200,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L200"
                  },
                  "summary": "True when the right-hand-side operand is stored transposed."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params_f32::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f32 activation",
                  "source": {
                    "line": 201,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L201"
                  },
                  "summary": "Output activation clamp range."
                },
                {
                  "description": "Right-hand-side operand storage format. `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` is currently supported only when `adj_x == false` and `adj_y == false`.",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params_f32::rhs_format",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rhs_format",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_weight_format_flt rhs_format",
                  "source": {
                    "line": 202,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L202"
                  },
                  "summary": "Right-hand-side operand storage format."
                }
              ],
              "name": "cmsis_nn_bmm_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_bmm_params_f32",
              "source": {
                "line": 197,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L197"
              },
              "summary": "Batched matrix multiplication parameters for float32 operators."
            },
            {
              "description": "Elementwise operator parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_ew_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_ew_params_f32::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f32 activation",
                  "source": {
                    "line": 213,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L213"
                  },
                  "summary": "Output activation clamp range."
                }
              ],
              "name": "cmsis_nn_ew_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_ew_params_f32",
              "source": {
                "line": 211,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L211"
              },
              "summary": "Elementwise operator parameters for float32 operators."
            },
            {
              "description": "Transpose parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_transpose_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Number of active dimensions in the permutation.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_params_f32::num_dims",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "num_dims",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t num_dims",
                  "source": {
                    "line": 221,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L221"
                  },
                  "summary": "Number of active dimensions in the permutation."
                },
                {
                  "description": "Permutation indices.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_params_f32::perm",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "perm",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t perm[4]",
                  "source": {
                    "line": 222,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L222"
                  },
                  "summary": "Permutation indices."
                },
                {
                  "description": "Layout convention used to interpret tensor dimensions.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_params_f32::layout",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "layout",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_tensor_layout layout",
                  "source": {
                    "line": 223,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L223"
                  },
                  "summary": "Layout convention used to interpret tensor dimensions."
                }
              ],
              "name": "cmsis_nn_transpose_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_transpose_params_f32",
              "source": {
                "line": 219,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L219"
              },
              "summary": "Transpose parameters for float32 operators."
            },
            {
              "description": "Singular value decomposition filter parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_svdf_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "SVDF rank.",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params_f32::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 231,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L231"
                  },
                  "summary": "SVDF rank."
                },
                {
                  "description": "Clamp range applied after the input projection.",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params_f32::input_activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f32 input_activation",
                  "source": {
                    "line": 232,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L232"
                  },
                  "summary": "Clamp range applied after the input projection."
                },
                {
                  "description": "Clamp range applied to the final output.",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params_f32::output_activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f32 output_activation",
                  "source": {
                    "line": 233,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L233"
                  },
                  "summary": "Clamp range applied to the final output."
                }
              ],
              "name": "cmsis_nn_svdf_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_svdf_params_f32",
              "source": {
                "line": 229,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L229"
              },
              "summary": "Singular value decomposition filter parameters for float32 operators."
            },
            {
              "description": "Read-only weights and bias metadata for one float32 LSTM gate.",
              "examples": [],
              "id": "cmsis_nn_lstm_gate_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Input-to-gate weight matrix.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate_f32::input_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float32_t * input_weights",
                  "source": {
                    "line": 241,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L241"
                  },
                  "summary": "Input-to-gate weight matrix."
                },
                {
                  "description": "Hidden-state-to-gate weight matrix.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate_f32::hidden_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float32_t * hidden_weights",
                  "source": {
                    "line": 242,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L242"
                  },
                  "summary": "Hidden-state-to-gate weight matrix."
                },
                {
                  "description": "Optional gate bias vector.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate_f32::bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float32_t * bias",
                  "source": {
                    "line": 243,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L243"
                  },
                  "summary": "Optional gate bias vector."
                },
                {
                  "description": "Gate activation selector.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate_f32::activation_type",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation_type",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_activation_type_flt activation_type",
                  "source": {
                    "line": 244,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L244"
                  },
                  "summary": "Gate activation selector."
                }
              ],
              "name": "cmsis_nn_lstm_gate_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_gate_f32",
              "source": {
                "line": 239,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L239"
              },
              "summary": "Read-only weights and bias metadata for one float32 LSTM gate."
            },
            {
              "description": "Parameters for a unidirectional float32 LSTM layer.",
              "examples": [],
              "id": "cmsis_nn_lstm_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Non-zero when input/output tensors are time-major.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::time_major",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_major",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_major",
                  "source": {
                    "line": 252,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L252"
                  },
                  "summary": "Non-zero when input/output tensors are time-major."
                },
                {
                  "description": "Batch size processed per invocation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::batch_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "batch_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t batch_size",
                  "source": {
                    "line": 253,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L253"
                  },
                  "summary": "Batch size processed per invocation."
                },
                {
                  "description": "Number of time steps processed per invocation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::time_steps",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_steps",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_steps",
                  "source": {
                    "line": 254,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L254"
                  },
                  "summary": "Number of time steps processed per invocation."
                },
                {
                  "description": "Input feature size per time step.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::input_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_size",
                  "source": {
                    "line": 255,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L255"
                  },
                  "summary": "Input feature size per time step."
                },
                {
                  "description": "Hidden-state size.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::hidden_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t hidden_size",
                  "source": {
                    "line": 256,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L256"
                  },
                  "summary": "Hidden-state size."
                },
                {
                  "description": "Optional cell-state clip value.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::cell_clip",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_clip",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float32_t cell_clip",
                  "source": {
                    "line": 257,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L257"
                  },
                  "summary": "Optional cell-state clip value."
                },
                {
                  "description": "Forget gate weights and activation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::forget_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "forget_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate_f32 forget_gate",
                  "source": {
                    "line": 259,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L259"
                  },
                  "summary": "Forget gate weights and activation."
                },
                {
                  "description": "Input gate weights and activation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::input_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate_f32 input_gate",
                  "source": {
                    "line": 260,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L260"
                  },
                  "summary": "Input gate weights and activation."
                },
                {
                  "description": "Cell-update gate weights and activation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::cell_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate_f32 cell_gate",
                  "source": {
                    "line": 261,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L261"
                  },
                  "summary": "Cell-update gate weights and activation."
                },
                {
                  "description": "Output gate weights and activation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f32::output_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate_f32 output_gate",
                  "source": {
                    "line": 262,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L262"
                  },
                  "summary": "Output gate weights and activation."
                }
              ],
              "name": "cmsis_nn_lstm_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_params_f32",
              "source": {
                "line": 250,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L250"
              },
              "summary": "Parameters for a unidirectional float32 LSTM layer."
            },
            {
              "description": "Scratch and mutable state buffers for a float32 LSTM invocation.",
              "examples": [],
              "id": "cmsis_nn_lstm_context_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Unused by the current implementation and may be NULL. Sized by `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, which reports 0; size from the query rather than hard-coding NULL if the buffer is arena-allocated.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context_f32::temp1",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "temp1",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float32_t * temp1",
                  "source": {
                    "line": 270,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L270"
                  },
                  "summary": "Unused by the current implementation and may be NULL."
                },
                {
                  "description": "Unused by the current implementation and may be NULL. Sized by `arm_lstm_unidirectional_f32_temp2_get_buffer_size()`.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context_f32::temp2",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "temp2",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float32_t * temp2",
                  "source": {
                    "line": 273,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L273"
                  },
                  "summary": "Unused by the current implementation and may be NULL."
                },
                {
                  "description": "Mutable cell-state buffer (in/out when streaming).",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context_f32::cell_state",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_state",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float32_t * cell_state",
                  "source": {
                    "line": 275,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L275"
                  },
                  "summary": "Mutable cell-state buffer (in/out when streaming)."
                },
                {
                  "description": "Optional in/out hidden state for streaming; NULL selects stateless. Streaming is NULL-gated, so zero-initialise the context (e.g. designated initialisers) to keep legacy 3-field callers stateless. Matches the quantized `cmsis_nn_lstm_context` contract.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context_f32::hidden_state",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_state",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float32_t * hidden_state",
                  "source": {
                    "line": 276,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L276"
                  },
                  "summary": "Optional in/out hidden state for streaming; NULL selects stateless."
                }
              ],
              "name": "cmsis_nn_lstm_context_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_context_f32",
              "source": {
                "line": 268,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L268"
              },
              "summary": "Scratch and mutable state buffers for a float32 LSTM invocation."
            },
            {
              "description": "Weights and biases for a single float32 GRU gate.\n\nThe activation (sigmoid for update/reset, tanh for candidate) is implied by the gate's role and is not stored here. The reset-after formulation keeps the input-projection bias and the recurrent-projection bias separate, because the reset gate multiplies the recurrent projection (including its bias) after the matmul.",
              "examples": [],
              "id": "cmsis_nn_gru_gate_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Input-to-gate weight matrix [hidden_size, input_size].",
                  "examples": [],
                  "id": "cmsis_nn_gru_gate_f32::input_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float32_t * input_weights",
                  "source": {
                    "line": 293,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L293"
                  },
                  "summary": "Input-to-gate weight matrix [hiddensize, inputsize]."
                },
                {
                  "description": "Hidden-to-gate weight matrix [hidden_size, hidden_size].",
                  "examples": [],
                  "id": "cmsis_nn_gru_gate_f32::hidden_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float32_t * hidden_weights",
                  "source": {
                    "line": 294,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L294"
                  },
                  "summary": "Hidden-to-gate weight matrix [hiddensize, hiddensize]."
                },
                {
                  "description": "Optional input-projection bias [hidden_size]. May be NULL.",
                  "examples": [],
                  "id": "cmsis_nn_gru_gate_f32::input_bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float32_t * input_bias",
                  "source": {
                    "line": 295,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L295"
                  },
                  "summary": "Optional input-projection bias [hiddensize]."
                },
                {
                  "description": "Optional recurrent-projection bias [hidden_size]. May be NULL.",
                  "examples": [],
                  "id": "cmsis_nn_gru_gate_f32::hidden_bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float32_t * hidden_bias",
                  "source": {
                    "line": 296,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L296"
                  },
                  "summary": "Optional recurrent-projection bias [hiddensize]."
                }
              ],
              "name": "cmsis_nn_gru_gate_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_gru_gate_f32",
              "source": {
                "line": 291,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L291"
              },
              "summary": "Weights and biases for a single float32 GRU gate."
            },
            {
              "description": "Parameters for a float32 unidirectional GRU invocation.\n\nGRU has three gates (update, reset, candidate) and, unlike LSTM, no cell state. The hidden state is the layer output.",
              "examples": [],
              "id": "cmsis_nn_gru_params_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Non-zero when input/output tensors are time-major.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f32::time_major",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_major",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_major",
                  "source": {
                    "line": 307,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L307"
                  },
                  "summary": "Non-zero when input/output tensors are time-major."
                },
                {
                  "description": "Batch size processed per invocation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f32::batch_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "batch_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t batch_size",
                  "source": {
                    "line": 308,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L308"
                  },
                  "summary": "Batch size processed per invocation."
                },
                {
                  "description": "Number of time steps processed per invocation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f32::time_steps",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_steps",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_steps",
                  "source": {
                    "line": 309,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L309"
                  },
                  "summary": "Number of time steps processed per invocation."
                },
                {
                  "description": "Input feature size per time step.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f32::input_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_size",
                  "source": {
                    "line": 310,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L310"
                  },
                  "summary": "Input feature size per time step."
                },
                {
                  "description": "Hidden-state size.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f32::hidden_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t hidden_size",
                  "source": {
                    "line": 311,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L311"
                  },
                  "summary": "Hidden-state size."
                },
                {
                  "description": "Non-zero: reset gate applied after the recurrent matmul (Keras/TFLite default).",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f32::reset_after",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "reset_after",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t reset_after",
                  "source": {
                    "line": 312,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L312"
                  },
                  "summary": "Non-zero: reset gate applied after the recurrent matmul (Keras/TFLite default)."
                },
                {
                  "description": "Update gate (z), sigmoid activation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f32::update_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "update_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_gru_gate_f32 update_gate",
                  "source": {
                    "line": 314,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L314"
                  },
                  "summary": "Update gate (z), sigmoid activation."
                },
                {
                  "description": "Reset gate (r), sigmoid activation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f32::reset_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "reset_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_gru_gate_f32 reset_gate",
                  "source": {
                    "line": 315,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L315"
                  },
                  "summary": "Reset gate (r), sigmoid activation."
                },
                {
                  "description": "Candidate/new gate (n), tanh activation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f32::candidate_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "candidate_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_gru_gate_f32 candidate_gate",
                  "source": {
                    "line": 316,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L316"
                  },
                  "summary": "Candidate/new gate (n), tanh activation."
                }
              ],
              "name": "cmsis_nn_gru_params_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_gru_params_f32",
              "source": {
                "line": 305,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L305"
              },
              "summary": "Parameters for a float32 unidirectional GRU invocation."
            },
            {
              "description": "Scratch buffers for a float32 GRU invocation.\n\n:::note\n`temp1` is sized by `arm_gru_unidirectional_f32_temp1_get_buffer_size()`: `hidden_size` elements when `reset_after == 0` (it holds the reset gate multiplied elementwise by the previous hidden state, r . h_prev). It is unused for the reset-after formulation and may be NULL there. There is no size field and no runtime enforcement: an undersized temp1 is written past on every build target.\n\n:::\n\n:::note\n`hidden_state` enables streaming state carry (`batch_size == 1`): when non-NULL it is read as the initial hidden state (seed to zero for a fresh sequence) and overwritten with the final hidden state on return. When NULL the state is zero-initialised and not written back.\n\n:::",
              "examples": [],
              "id": "cmsis_nn_gru_context_f32",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Scratch required when reset_after == 0; sized by `arm_gru_unidirectional_f32_temp1_get_buffer_size()`.",
                  "examples": [],
                  "id": "cmsis_nn_gru_context_f32::temp1",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "temp1",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float32_t * temp1",
                  "source": {
                    "line": 335,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L335"
                  },
                  "summary": "Scratch required when resetafter == 0; sized by armgruunidirectionalf32temp1getbuffersize()."
                },
                {
                  "description": "Optional in/out persistent hidden state [hidden_size] for streaming (batch_size == 1).",
                  "examples": [],
                  "id": "cmsis_nn_gru_context_f32::hidden_state",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_state",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float32_t * hidden_state",
                  "source": {
                    "line": 338,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L338"
                  },
                  "summary": "Optional in/out persistent hidden state [hiddensize] for streaming (batchsize == 1)."
                }
              ],
              "name": "cmsis_nn_gru_context_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_gru_context_f32",
              "source": {
                "line": 333,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L333"
              },
              "summary": "Scratch buffers for a float32 GRU invocation."
            },
            {
              "description": "Activation clamp range for floating-point operators.",
              "examples": [],
              "id": "cmsis_nn_activation_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Minimum value used to clamp the result.",
                  "examples": [],
                  "id": "cmsis_nn_activation_f16::min",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "min",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float16_t min",
                  "source": {
                    "line": 352,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L352"
                  },
                  "summary": "Minimum value used to clamp the result."
                },
                {
                  "description": "Maximum value used to clamp the result.",
                  "examples": [],
                  "id": "cmsis_nn_activation_f16::max",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "max",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float16_t max",
                  "source": {
                    "line": 353,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L353"
                  },
                  "summary": "Maximum value used to clamp the result."
                }
              ],
              "name": "cmsis_nn_activation_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_activation_f16",
              "source": {
                "line": 350,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L350"
              },
              "summary": "Activation clamp range for floating-point operators."
            },
            {
              "description": "Convolution parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_conv_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Spatial stride.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f16::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 361,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L361"
                  },
                  "summary": "Spatial stride."
                },
                {
                  "description": "Spatial zero-padding.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f16::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 362,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L362"
                  },
                  "summary": "Spatial zero-padding."
                },
                {
                  "description": "Spatial dilation.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f16::dilation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "dilation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile dilation",
                  "source": {
                    "line": 363,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L363"
                  },
                  "summary": "Spatial dilation."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f16::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f16 activation",
                  "source": {
                    "line": 364,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L364"
                  },
                  "summary": "Output activation clamp range."
                },
                {
                  "description": "Filter storage format.",
                  "examples": [],
                  "id": "cmsis_nn_conv_params_f16::weight_format",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "weight_format",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_weight_format_flt weight_format",
                  "source": {
                    "line": 365,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L365"
                  },
                  "summary": "Filter storage format."
                }
              ],
              "name": "cmsis_nn_conv_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_conv_params_f16",
              "source": {
                "line": 359,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L359"
              },
              "summary": "Convolution parameters for float32 operators."
            },
            {
              "description": "Transpose convolution parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_transpose_conv_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Spatial stride.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f16::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 373,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L373"
                  },
                  "summary": "Spatial stride."
                },
                {
                  "description": "Spatial zero-padding.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f16::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 374,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L374"
                  },
                  "summary": "Spatial zero-padding."
                },
                {
                  "description": "Output padding adjustment for transpose convolution.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f16::padding_offsets",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding_offsets",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding_offsets",
                  "source": {
                    "line": 375,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L375"
                  },
                  "summary": "Output padding adjustment for transpose convolution."
                },
                {
                  "description": "Spatial dilation.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f16::dilation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "dilation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile dilation",
                  "source": {
                    "line": 376,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L376"
                  },
                  "summary": "Spatial dilation."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_conv_params_f16::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f16 activation",
                  "source": {
                    "line": 377,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L377"
                  },
                  "summary": "Output activation clamp range."
                }
              ],
              "name": "cmsis_nn_transpose_conv_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_transpose_conv_params_f16",
              "source": {
                "line": 371,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L371"
              },
              "summary": "Transpose convolution parameters for float32 operators."
            },
            {
              "description": "Depthwise convolution parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_dw_conv_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Channel multiplier. `ch_mult * in_ch = out_ch`.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f16::ch_mult",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "ch_mult",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t ch_mult",
                  "source": {
                    "line": 385,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L385"
                  },
                  "summary": "Channel multiplier."
                },
                {
                  "description": "Spatial stride.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f16::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 386,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L386"
                  },
                  "summary": "Spatial stride."
                },
                {
                  "description": "Spatial zero-padding.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f16::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 387,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L387"
                  },
                  "summary": "Spatial zero-padding."
                },
                {
                  "description": "Spatial dilation.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f16::dilation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "dilation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile dilation",
                  "source": {
                    "line": 388,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L388"
                  },
                  "summary": "Spatial dilation."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_dw_conv_params_f16::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f16 activation",
                  "source": {
                    "line": 389,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L389"
                  },
                  "summary": "Output activation clamp range."
                }
              ],
              "name": "cmsis_nn_dw_conv_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_dw_conv_params_f16",
              "source": {
                "line": 383,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L383"
              },
              "summary": "Depthwise convolution parameters for float32 operators."
            },
            {
              "description": "Pooling parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_pool_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Spatial stride.",
                  "examples": [],
                  "id": "cmsis_nn_pool_params_f16::stride",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "stride",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile stride",
                  "source": {
                    "line": 397,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L397"
                  },
                  "summary": "Spatial stride."
                },
                {
                  "description": "Spatial zero-padding.",
                  "examples": [],
                  "id": "cmsis_nn_pool_params_f16::padding",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "padding",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_tile padding",
                  "source": {
                    "line": 398,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L398"
                  },
                  "summary": "Spatial zero-padding."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_pool_params_f16::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f16 activation",
                  "source": {
                    "line": 399,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L399"
                  },
                  "summary": "Output activation clamp range."
                }
              ],
              "name": "cmsis_nn_pool_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_pool_params_f16",
              "source": {
                "line": 395,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L395"
              },
              "summary": "Pooling parameters for float32 operators."
            },
            {
              "description": "Fully connected layer parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_fc_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_fc_params_f16::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f16 activation",
                  "source": {
                    "line": 407,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L407"
                  },
                  "summary": "Output activation clamp range."
                },
                {
                  "description": "Weight storage format.",
                  "examples": [],
                  "id": "cmsis_nn_fc_params_f16::weight_format",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "weight_format",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_weight_format_flt weight_format",
                  "source": {
                    "line": 408,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L408"
                  },
                  "summary": "Weight storage format."
                }
              ],
              "name": "cmsis_nn_fc_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_fc_params_f16",
              "source": {
                "line": 405,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L405"
              },
              "summary": "Fully connected layer parameters for float32 operators."
            },
            {
              "description": "Batched matrix multiplication parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_bmm_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "True when the left-hand-side operand is stored transposed.",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params_f16::adj_x",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "adj_x",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const bool adj_x",
                  "source": {
                    "line": 416,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L416"
                  },
                  "summary": "True when the left-hand-side operand is stored transposed."
                },
                {
                  "description": "True when the right-hand-side operand is stored transposed.",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params_f16::adj_y",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "adj_y",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const bool adj_y",
                  "source": {
                    "line": 417,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L417"
                  },
                  "summary": "True when the right-hand-side operand is stored transposed."
                },
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params_f16::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f16 activation",
                  "source": {
                    "line": 418,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L418"
                  },
                  "summary": "Output activation clamp range."
                },
                {
                  "description": "Right-hand-side operand storage format. `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` is currently supported only when `adj_x == false` and `adj_y == false`.",
                  "examples": [],
                  "id": "cmsis_nn_bmm_params_f16::rhs_format",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rhs_format",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_weight_format_flt rhs_format",
                  "source": {
                    "line": 419,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L419"
                  },
                  "summary": "Right-hand-side operand storage format."
                }
              ],
              "name": "cmsis_nn_bmm_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_bmm_params_f16",
              "source": {
                "line": 414,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L414"
              },
              "summary": "Batched matrix multiplication parameters for float32 operators."
            },
            {
              "description": "Elementwise operator parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_ew_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Output activation clamp range.",
                  "examples": [],
                  "id": "cmsis_nn_ew_params_f16::activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f16 activation",
                  "source": {
                    "line": 430,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L430"
                  },
                  "summary": "Output activation clamp range."
                }
              ],
              "name": "cmsis_nn_ew_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_ew_params_f16",
              "source": {
                "line": 428,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L428"
              },
              "summary": "Elementwise operator parameters for float32 operators."
            },
            {
              "description": "Transpose parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_transpose_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Number of active dimensions in the permutation.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_params_f16::num_dims",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "num_dims",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t num_dims",
                  "source": {
                    "line": 438,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L438"
                  },
                  "summary": "Number of active dimensions in the permutation."
                },
                {
                  "description": "Permutation indices.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_params_f16::perm",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "perm",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t perm[4]",
                  "source": {
                    "line": 439,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L439"
                  },
                  "summary": "Permutation indices."
                },
                {
                  "description": "Layout convention used to interpret tensor dimensions.",
                  "examples": [],
                  "id": "cmsis_nn_transpose_params_f16::layout",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "layout",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_tensor_layout layout",
                  "source": {
                    "line": 440,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L440"
                  },
                  "summary": "Layout convention used to interpret tensor dimensions."
                }
              ],
              "name": "cmsis_nn_transpose_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_transpose_params_f16",
              "source": {
                "line": 436,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L436"
              },
              "summary": "Transpose parameters for float32 operators."
            },
            {
              "description": "Singular value decomposition filter parameters for float32 operators.",
              "examples": [],
              "id": "cmsis_nn_svdf_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "SVDF rank.",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params_f16::rank",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "rank",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t rank",
                  "source": {
                    "line": 448,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L448"
                  },
                  "summary": "SVDF rank."
                },
                {
                  "description": "Clamp range applied after the input projection.",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params_f16::input_activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f16 input_activation",
                  "source": {
                    "line": 449,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L449"
                  },
                  "summary": "Clamp range applied after the input projection."
                },
                {
                  "description": "Clamp range applied to the final output.",
                  "examples": [],
                  "id": "cmsis_nn_svdf_params_f16::output_activation",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_activation",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_activation_f16 output_activation",
                  "source": {
                    "line": 450,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L450"
                  },
                  "summary": "Clamp range applied to the final output."
                }
              ],
              "name": "cmsis_nn_svdf_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_svdf_params_f16",
              "source": {
                "line": 446,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L446"
              },
              "summary": "Singular value decomposition filter parameters for float32 operators."
            },
            {
              "description": "Read-only weights and bias metadata for one float32 LSTM gate.",
              "examples": [],
              "id": "cmsis_nn_lstm_gate_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Input-to-gate weight matrix.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate_f16::input_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float16_t * input_weights",
                  "source": {
                    "line": 458,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L458"
                  },
                  "summary": "Input-to-gate weight matrix."
                },
                {
                  "description": "Hidden-state-to-gate weight matrix.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate_f16::hidden_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float16_t * hidden_weights",
                  "source": {
                    "line": 459,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L459"
                  },
                  "summary": "Hidden-state-to-gate weight matrix."
                },
                {
                  "description": "Optional gate bias vector.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate_f16::bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float16_t * bias",
                  "source": {
                    "line": 460,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L460"
                  },
                  "summary": "Optional gate bias vector."
                },
                {
                  "description": "Gate activation selector.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_gate_f16::activation_type",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "activation_type",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "arm_nn_activation_type_flt activation_type",
                  "source": {
                    "line": 461,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L461"
                  },
                  "summary": "Gate activation selector."
                }
              ],
              "name": "cmsis_nn_lstm_gate_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_gate_f16",
              "source": {
                "line": 456,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L456"
              },
              "summary": "Read-only weights and bias metadata for one float32 LSTM gate."
            },
            {
              "description": "Parameters for a unidirectional float32 LSTM layer.",
              "examples": [],
              "id": "cmsis_nn_lstm_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Non-zero when input/output tensors are time-major.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::time_major",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_major",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_major",
                  "source": {
                    "line": 469,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L469"
                  },
                  "summary": "Non-zero when input/output tensors are time-major."
                },
                {
                  "description": "Batch size processed per invocation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::batch_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "batch_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t batch_size",
                  "source": {
                    "line": 470,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L470"
                  },
                  "summary": "Batch size processed per invocation."
                },
                {
                  "description": "Number of time steps processed per invocation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::time_steps",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_steps",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_steps",
                  "source": {
                    "line": 471,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L471"
                  },
                  "summary": "Number of time steps processed per invocation."
                },
                {
                  "description": "Input feature size per time step.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::input_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_size",
                  "source": {
                    "line": 472,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L472"
                  },
                  "summary": "Input feature size per time step."
                },
                {
                  "description": "Hidden-state size.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::hidden_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t hidden_size",
                  "source": {
                    "line": 473,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L473"
                  },
                  "summary": "Hidden-state size."
                },
                {
                  "description": "Optional cell-state clip value.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::cell_clip",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_clip",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float16_t cell_clip",
                  "source": {
                    "line": 474,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L474"
                  },
                  "summary": "Optional cell-state clip value."
                },
                {
                  "description": "Forget gate weights and activation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::forget_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "forget_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate_f16 forget_gate",
                  "source": {
                    "line": 476,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L476"
                  },
                  "summary": "Forget gate weights and activation."
                },
                {
                  "description": "Input gate weights and activation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::input_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate_f16 input_gate",
                  "source": {
                    "line": 477,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L477"
                  },
                  "summary": "Input gate weights and activation."
                },
                {
                  "description": "Cell-update gate weights and activation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::cell_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate_f16 cell_gate",
                  "source": {
                    "line": 478,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L478"
                  },
                  "summary": "Cell-update gate weights and activation."
                },
                {
                  "description": "Output gate weights and activation.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_params_f16::output_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "output_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_lstm_gate_f16 output_gate",
                  "source": {
                    "line": 479,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L479"
                  },
                  "summary": "Output gate weights and activation."
                }
              ],
              "name": "cmsis_nn_lstm_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_params_f16",
              "source": {
                "line": 467,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L467"
              },
              "summary": "Parameters for a unidirectional float32 LSTM layer."
            },
            {
              "description": "Scratch and mutable state buffers for a float32 LSTM invocation.",
              "examples": [],
              "id": "cmsis_nn_lstm_context_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Unused by the current implementation and may be NULL. Sized by `arm_lstm_unidirectional_f16_temp1_get_buffer_size()`, which reports 0; size from the query rather than hard-coding NULL if the buffer is arena-allocated.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context_f16::temp1",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "temp1",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float16_t * temp1",
                  "source": {
                    "line": 487,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L487"
                  },
                  "summary": "Unused by the current implementation and may be NULL."
                },
                {
                  "description": "Unused by the current implementation and may be NULL. Sized by `arm_lstm_unidirectional_f16_temp2_get_buffer_size()`.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context_f16::temp2",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "temp2",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float16_t * temp2",
                  "source": {
                    "line": 490,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L490"
                  },
                  "summary": "Unused by the current implementation and may be NULL."
                },
                {
                  "description": "Mutable cell-state buffer (in/out when streaming).",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context_f16::cell_state",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "cell_state",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float16_t * cell_state",
                  "source": {
                    "line": 492,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L492"
                  },
                  "summary": "Mutable cell-state buffer (in/out when streaming)."
                },
                {
                  "description": "Optional in/out hidden state for streaming; NULL selects stateless. Streaming is NULL-gated, so zero-initialise the context (e.g. designated initialisers) to keep legacy 3-field callers stateless. Matches the quantized `cmsis_nn_lstm_context` contract.",
                  "examples": [],
                  "id": "cmsis_nn_lstm_context_f16::hidden_state",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_state",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float16_t * hidden_state",
                  "source": {
                    "line": 493,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L493"
                  },
                  "summary": "Optional in/out hidden state for streaming; NULL selects stateless."
                }
              ],
              "name": "cmsis_nn_lstm_context_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_lstm_context_f16",
              "source": {
                "line": 485,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L485"
              },
              "summary": "Scratch and mutable state buffers for a float32 LSTM invocation."
            },
            {
              "description": "Weights and biases for a single float16 GRU gate.\n\nThe activation (sigmoid for update/reset, tanh for candidate) is implied by the gate's role and is not stored here. The reset-after formulation keeps the input-projection bias and the recurrent-projection bias separate, because the reset gate multiplies the recurrent projection (including its bias) after the matmul.",
              "examples": [],
              "id": "cmsis_nn_gru_gate_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Input-to-gate weight matrix [hidden_size, input_size].",
                  "examples": [],
                  "id": "cmsis_nn_gru_gate_f16::input_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float16_t * input_weights",
                  "source": {
                    "line": 510,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L510"
                  },
                  "summary": "Input-to-gate weight matrix [hiddensize, inputsize]."
                },
                {
                  "description": "Hidden-to-gate weight matrix [hidden_size, hidden_size].",
                  "examples": [],
                  "id": "cmsis_nn_gru_gate_f16::hidden_weights",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_weights",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float16_t * hidden_weights",
                  "source": {
                    "line": 511,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L511"
                  },
                  "summary": "Hidden-to-gate weight matrix [hiddensize, hiddensize]."
                },
                {
                  "description": "Optional input-projection bias [hidden_size]. May be NULL.",
                  "examples": [],
                  "id": "cmsis_nn_gru_gate_f16::input_bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float16_t * input_bias",
                  "source": {
                    "line": 512,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L512"
                  },
                  "summary": "Optional input-projection bias [hiddensize]."
                },
                {
                  "description": "Optional recurrent-projection bias [hidden_size]. May be NULL.",
                  "examples": [],
                  "id": "cmsis_nn_gru_gate_f16::hidden_bias",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_bias",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const float16_t * hidden_bias",
                  "source": {
                    "line": 513,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L513"
                  },
                  "summary": "Optional recurrent-projection bias [hiddensize]."
                }
              ],
              "name": "cmsis_nn_gru_gate_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_gru_gate_f16",
              "source": {
                "line": 508,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L508"
              },
              "summary": "Weights and biases for a single float16 GRU gate."
            },
            {
              "description": "Parameters for a float16 unidirectional GRU invocation.\n\nGRU has three gates (update, reset, candidate) and, unlike LSTM, no cell state. The hidden state is the layer output.",
              "examples": [],
              "id": "cmsis_nn_gru_params_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Non-zero when input/output tensors are time-major.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f16::time_major",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_major",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_major",
                  "source": {
                    "line": 524,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L524"
                  },
                  "summary": "Non-zero when input/output tensors are time-major."
                },
                {
                  "description": "Batch size processed per invocation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f16::batch_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "batch_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t batch_size",
                  "source": {
                    "line": 525,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L525"
                  },
                  "summary": "Batch size processed per invocation."
                },
                {
                  "description": "Number of time steps processed per invocation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f16::time_steps",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "time_steps",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t time_steps",
                  "source": {
                    "line": 526,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L526"
                  },
                  "summary": "Number of time steps processed per invocation."
                },
                {
                  "description": "Input feature size per time step.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f16::input_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "input_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t input_size",
                  "source": {
                    "line": 527,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L527"
                  },
                  "summary": "Input feature size per time step."
                },
                {
                  "description": "Hidden-state size.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f16::hidden_size",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_size",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t hidden_size",
                  "source": {
                    "line": 528,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L528"
                  },
                  "summary": "Hidden-state size."
                },
                {
                  "description": "Non-zero: reset gate applied after the recurrent matmul (Keras/TFLite default).",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f16::reset_after",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "reset_after",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "int32_t reset_after",
                  "source": {
                    "line": 529,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L529"
                  },
                  "summary": "Non-zero: reset gate applied after the recurrent matmul (Keras/TFLite default)."
                },
                {
                  "description": "Update gate (z), sigmoid activation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f16::update_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "update_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_gru_gate_f16 update_gate",
                  "source": {
                    "line": 531,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L531"
                  },
                  "summary": "Update gate (z), sigmoid activation."
                },
                {
                  "description": "Reset gate (r), sigmoid activation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f16::reset_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "reset_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_gru_gate_f16 reset_gate",
                  "source": {
                    "line": 532,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L532"
                  },
                  "summary": "Reset gate (r), sigmoid activation."
                },
                {
                  "description": "Candidate/new gate (n), tanh activation.",
                  "examples": [],
                  "id": "cmsis_nn_gru_params_f16::candidate_gate",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "candidate_gate",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "cmsis_nn_gru_gate_f16 candidate_gate",
                  "source": {
                    "line": 533,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L533"
                  },
                  "summary": "Candidate/new gate (n), tanh activation."
                }
              ],
              "name": "cmsis_nn_gru_params_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_gru_params_f16",
              "source": {
                "line": 522,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L522"
              },
              "summary": "Parameters for a float16 unidirectional GRU invocation."
            },
            {
              "description": "Scratch buffers for a float16 GRU invocation.\n\n:::note\n`temp1` is sized by `arm_gru_unidirectional_f16_temp1_get_buffer_size()`: `hidden_size` elements when `reset_after == 0` (it holds the reset gate multiplied elementwise by the previous hidden state, r . h_prev). It is unused for the reset-after formulation and may be NULL there. There is no size field and no runtime enforcement: an undersized temp1 is written past on every build target.\n\n:::\n\n:::note\n`hidden_state` enables streaming state carry (`batch_size == 1`): when non-NULL it is read as the initial hidden state (seed to zero for a fresh sequence) and overwritten with the final hidden state on return. When NULL the state is zero-initialised and not written back.\n\n:::",
              "examples": [],
              "id": "cmsis_nn_gru_context_f16",
              "kind": "struct",
              "language": "c",
              "members": [
                {
                  "description": "Scratch required when reset_after == 0; sized by `arm_gru_unidirectional_f16_temp1_get_buffer_size()`.",
                  "examples": [],
                  "id": "cmsis_nn_gru_context_f16::temp1",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "temp1",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float16_t * temp1",
                  "source": {
                    "line": 552,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L552"
                  },
                  "summary": "Scratch required when resetafter == 0; sized by armgruunidirectionalf16temp1getbuffersize()."
                },
                {
                  "description": "Optional in/out persistent hidden state [hidden_size] for streaming (batch_size == 1).",
                  "examples": [],
                  "id": "cmsis_nn_gru_context_f16::hidden_state",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "hidden_state",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "float16_t * hidden_state",
                  "source": {
                    "line": 555,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L555"
                  },
                  "summary": "Optional in/out persistent hidden state [hiddensize] for streaming (batchsize == 1)."
                }
              ],
              "name": "cmsis_nn_gru_context_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "struct cmsis_nn_gru_context_f16",
              "source": {
                "line": 550,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L550"
              },
              "summary": "Scratch buffers for a float16 GRU invocation."
            },
            {
              "description": "Tensor layout selector for floating-point APIs.\n\nFloat public APIs currently accept NHWC layout only.",
              "examples": [],
              "id": "arm_nn_tensor_layout",
              "kind": "enum",
              "language": "c",
              "members": [
                {
                  "description": "Tensor dimensions are ordered as [N, H, W, C].",
                  "examples": [],
                  "id": "ARM_NN_LAYOUT_NHWC",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_LAYOUT_NHWC",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_LAYOUT_NHWC = 0",
                  "source": {
                    "line": 67,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L67"
                  },
                  "summary": "Tensor dimensions are ordered as [N, H, W, C]."
                }
              ],
              "name": "arm_nn_tensor_layout",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "enum arm_nn_tensor_layout",
              "source": {
                "line": 67,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L67"
              },
              "summary": "Tensor layout selector for floating-point APIs."
            },
            {
              "description": "Enum for specifying activation function types",
              "examples": [],
              "id": "arm_nn_activation_type",
              "kind": "enum",
              "language": "c",
              "members": [
                {
                  "description": "Sigmoid activation function",
                  "examples": [],
                  "id": "ARM_SIGMOID",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_SIGMOID",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_SIGMOID = 0",
                  "source": {
                    "line": 78,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L78"
                  },
                  "summary": "Sigmoid activation function"
                },
                {
                  "description": "Tanh activation function",
                  "examples": [],
                  "id": "ARM_TANH",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_TANH",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_TANH = 1",
                  "source": {
                    "line": 78,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L78"
                  },
                  "summary": "Tanh activation function"
                }
              ],
              "name": "arm_nn_activation_type",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "enum arm_nn_activation_type",
              "source": {
                "line": 78,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L78"
              },
              "summary": "Enum for specifying activation function types"
            },
            {
              "description": "Activation selector for floating-point operator APIs.\n\nNumeric values intentionally live in a dedicated floating-point range to avoid overlap with the legacy integer public activation enum.",
              "examples": [],
              "id": "arm_nn_activation_type_flt",
              "kind": "enum",
              "language": "c",
              "members": [
                {
                  "description": "Identity activation function.",
                  "examples": [],
                  "id": "ARM_NN_FLT_ACT_NONE",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_FLT_ACT_NONE",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_FLT_ACT_NONE = 32",
                  "source": {
                    "line": 78,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L78"
                  },
                  "summary": "Identity activation function."
                },
                {
                  "description": "Sigmoid activation function.",
                  "examples": [],
                  "id": "ARM_NN_FLT_ACT_SIGMOID",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_FLT_ACT_SIGMOID",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_FLT_ACT_SIGMOID = 33",
                  "source": {
                    "line": 78,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L78"
                  },
                  "summary": "Sigmoid activation function."
                },
                {
                  "description": "Hyperbolic tangent activation function.",
                  "examples": [],
                  "id": "ARM_NN_FLT_ACT_TANH",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_FLT_ACT_TANH",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_FLT_ACT_TANH = 34",
                  "source": {
                    "line": 78,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L78"
                  },
                  "summary": "Hyperbolic tangent activation function."
                },
                {
                  "description": "ReLU activation function.",
                  "examples": [],
                  "id": "ARM_NN_FLT_ACT_RELU",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_FLT_ACT_RELU",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_FLT_ACT_RELU = 35",
                  "source": {
                    "line": 78,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L78"
                  },
                  "summary": "ReLU activation function."
                },
                {
                  "description": "ReLU6 activation function.",
                  "examples": [],
                  "id": "ARM_NN_FLT_ACT_RELU6",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_FLT_ACT_RELU6",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_FLT_ACT_RELU6 = 36",
                  "source": {
                    "line": 78,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L78"
                  },
                  "summary": "ReLU6 activation function."
                },
                {
                  "description": "Hard-swish activation function.",
                  "examples": [],
                  "id": "ARM_NN_FLT_ACT_HARDSWISH",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_FLT_ACT_HARDSWISH",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_FLT_ACT_HARDSWISH = 37",
                  "source": {
                    "line": 78,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L78"
                  },
                  "summary": "Hard-swish activation function."
                },
                {
                  "description": "Leaky ReLU activation function.",
                  "examples": [],
                  "id": "ARM_NN_FLT_ACT_LEAKY_RELU",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_FLT_ACT_LEAKY_RELU",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_FLT_ACT_LEAKY_RELU = 38",
                  "source": {
                    "line": 78,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L78"
                  },
                  "summary": "Leaky ReLU activation function."
                }
              ],
              "name": "arm_nn_activation_type_flt",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "enum arm_nn_activation_type_flt",
              "source": {
                "line": 78,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L78"
              },
              "summary": "Activation selector for floating-point operator APIs."
            },
            {
              "description": "Enum for specifying comparison operator",
              "examples": [],
              "id": "arm_nn_compare_operation",
              "kind": "enum",
              "language": "c",
              "members": [
                {
                  "description": "Returns 1 if lhs == rhs else 0",
                  "examples": [],
                  "id": "ARM_COMPARE_EQUAL",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_COMPARE_EQUAL",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_COMPARE_EQUAL = 0",
                  "source": {
                    "line": 85,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L85"
                  },
                  "summary": "Returns 1 if lhs == rhs else 0"
                },
                {
                  "description": "Returns 1 if lhs != rhs else 0",
                  "examples": [],
                  "id": "ARM_COMPARE_NOT_EQUAL",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_COMPARE_NOT_EQUAL",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_COMPARE_NOT_EQUAL = 1",
                  "source": {
                    "line": 85,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L85"
                  },
                  "summary": "Returns 1 if lhs != rhs else 0"
                },
                {
                  "description": "Returns 1 if lhs > rhs else 0",
                  "examples": [],
                  "id": "ARM_COMPARE_GREATER",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_COMPARE_GREATER",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_COMPARE_GREATER = 2",
                  "source": {
                    "line": 85,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L85"
                  },
                  "summary": "Returns 1 if lhs > rhs else 0"
                },
                {
                  "description": "Returns 1 if lhs >= rhs else 0",
                  "examples": [],
                  "id": "ARM_COMPARE_GREATER_EQUAL",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_COMPARE_GREATER_EQUAL",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_COMPARE_GREATER_EQUAL = 3",
                  "source": {
                    "line": 85,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L85"
                  },
                  "summary": "Returns 1 if lhs >= rhs else 0"
                },
                {
                  "description": "Returns 1 if lhs < rhs else 0",
                  "examples": [],
                  "id": "ARM_COMPARE_LESS",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_COMPARE_LESS",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_COMPARE_LESS = 4",
                  "source": {
                    "line": 85,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L85"
                  },
                  "summary": "Returns 1 if lhs < rhs else 0"
                },
                {
                  "description": "Returns 1 if lhs <= rhs else 0",
                  "examples": [],
                  "id": "ARM_COMPARE_LESS_EQUAL",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_COMPARE_LESS_EQUAL",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_COMPARE_LESS_EQUAL = 5",
                  "source": {
                    "line": 85,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L85"
                  },
                  "summary": "Returns 1 if lhs <= rhs else 0"
                }
              ],
              "name": "arm_nn_compare_operation",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "enum arm_nn_compare_operation",
              "source": {
                "line": 85,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L85"
              },
              "summary": "Enum for specifying comparison operator"
            },
            {
              "description": "Function return codes",
              "examples": [],
              "id": "arm_cmsis_nn_status",
              "kind": "enum",
              "language": "c",
              "members": [
                {
                  "description": "No error",
                  "examples": [],
                  "id": "ARM_CMSIS_NN_SUCCESS",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CMSIS_NN_SUCCESS",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_CMSIS_NN_SUCCESS = 0",
                  "source": {
                    "line": 96,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L96"
                  },
                  "summary": "No error"
                },
                {
                  "description": "One or more arguments are incorrect",
                  "examples": [],
                  "id": "ARM_CMSIS_NN_ARG_ERROR",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CMSIS_NN_ARG_ERROR",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_CMSIS_NN_ARG_ERROR = -1",
                  "source": {
                    "line": 96,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L96"
                  },
                  "summary": "One or more arguments are incorrect"
                },
                {
                  "description": "No implementation available",
                  "examples": [],
                  "id": "ARM_CMSIS_NN_NO_IMPL_ERROR",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CMSIS_NN_NO_IMPL_ERROR",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_CMSIS_NN_NO_IMPL_ERROR = -2",
                  "source": {
                    "line": 96,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L96"
                  },
                  "summary": "No implementation available"
                },
                {
                  "description": "Logical error",
                  "examples": [],
                  "id": "ARM_CMSIS_NN_FAILURE",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CMSIS_NN_FAILURE",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_CMSIS_NN_FAILURE = -3",
                  "source": {
                    "line": 96,
                    "path": "Include/arm_nn_types.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L96"
                  },
                  "summary": "Logical error"
                }
              ],
              "name": "arm_cmsis_nn_status",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "enum arm_cmsis_nn_status",
              "source": {
                "line": 96,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L96"
              },
              "summary": "Function return codes"
            },
            {
              "description": "Depthwise kernel storage layout selector for floating-point kernels.\n\nPublic float depthwise entry points currently use KC storage (`[k][c]`).",
              "examples": [],
              "id": "arm_nn_dw_kernel_layout_f32",
              "kind": "enum",
              "language": "c",
              "members": [
                {
                  "description": "Depthwise kernel stored as `[kernel][channel]`.",
                  "examples": [],
                  "id": "ARM_NN_DW_KERNEL_KC",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_DW_KERNEL_KC",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_DW_KERNEL_KC = 0",
                  "source": {
                    "line": 98,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L98"
                  },
                  "summary": "Depthwise kernel stored as [kernel][channel]."
                },
                {
                  "description": "Depthwise kernel stored as `[channel][kernel]`.",
                  "examples": [],
                  "id": "ARM_NN_DW_KERNEL_CK",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_DW_KERNEL_CK",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_DW_KERNEL_CK = 1",
                  "source": {
                    "line": 98,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L98"
                  },
                  "summary": "Depthwise kernel stored as [channel][kernel]."
                }
              ],
              "name": "arm_nn_dw_kernel_layout_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "enum arm_nn_dw_kernel_layout_f32",
              "source": {
                "line": 98,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L98"
              },
              "summary": "Depthwise kernel storage layout selector for floating-point kernels."
            },
            {
              "description": "Weight storage format selector for floating-point operators.\n\nThis enum allows frameworks to describe whether weights are provided in the standard public operator layout or in a backend-specific packed layout.\n\n`ARM_NN_WEIGHT_FORMAT_NT_N_PACKED` matches the packed RHS layout consumed by `arm_nn_mat_mult_nt_n_packed_f16/f32`.\n\nThe packed NTxN layout exists because MVE kernels typically perform best when output-channel blocks can be loaded contiguously. With the standard `NT x T` formulation, vectorizing over output channels tends to require gather-load accesses to the RHS, which is less efficient than a packed non-transposed RHS layout.\n\nFor operators that support `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, supplying offline-repacked constant weights in this layout is therefore generally the preferred way to achieve the best MVE performance.",
              "examples": [],
              "id": "arm_nn_weight_format_flt",
              "kind": "enum",
              "language": "c",
              "members": [
                {
                  "description": "Standard public operator layout.",
                  "examples": [],
                  "id": "ARM_NN_WEIGHT_FORMAT_STANDARD",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_WEIGHT_FORMAT_STANDARD",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_WEIGHT_FORMAT_STANDARD = 0",
                  "source": {
                    "line": 122,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L122"
                  },
                  "summary": "Standard public operator layout."
                },
                {
                  "description": "Packed `[K][N-block]` layout for NTxN matmul helpers.",
                  "examples": [],
                  "id": "ARM_NN_WEIGHT_FORMAT_NT_N_PACKED",
                  "kind": "constant",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_WEIGHT_FORMAT_NT_N_PACKED",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "ARM_NN_WEIGHT_FORMAT_NT_N_PACKED = 1",
                  "source": {
                    "line": 122,
                    "path": "Include/arm_nn_types_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L122"
                  },
                  "summary": "Packed [K][N-block] layout for NTxN matmul helpers."
                }
              ],
              "name": "arm_nn_weight_format_flt",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "enum arm_nn_weight_format_flt",
              "source": {
                "line": 122,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L122"
              },
              "summary": "Weight storage format selector for floating-point operators."
            },
            {
              "description": "",
              "examples": [],
              "id": "arm_nn_dw_kernel_layout_f16",
              "kind": "type",
              "language": "c",
              "members": [],
              "name": "arm_nn_dw_kernel_layout_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "typedef arm_nn_dw_kernel_layout_f32 arm_nn_dw_kernel_layout_f16",
              "source": {
                "line": 345,
                "path": "Include/arm_nn_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types_flt.h#L345"
              },
              "summary": ""
            }
          ]
        },
        {
          "description": "Elementwise add and multiplication functions.",
          "name": "Elementwise Functions",
          "path": "heliaCORE.groupElementwise",
          "submodules": [],
          "summary": "Elementwise add and multiplication functions.",
          "symbols": [
            {
              "description": "Elementwise add with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
              "examples": [],
              "id": "arm_elementwise_add_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_add_f32",
              "params": [
                {
                  "description": "Pointer to the first input vector.",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the second input vector.",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float32_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float32_t"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_add_f32(\n    const float32_t *input_1_vect,\n    const float32_t *input_2_vect,\n    float32_t *output,\n    float32_t out_activation_min,\n    float32_t out_activation_max,\n    int32_t block_size\n)",
              "source": {
                "line": 714,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L714"
              },
              "summary": "Elementwise add with optional output clamp."
            },
            {
              "description": "Elementwise subtract with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
              "examples": [],
              "id": "arm_elementwise_sub_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_sub_f32",
              "params": [
                {
                  "description": "Pointer to the first input vector (minuend).",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the second input vector (subtrahend).",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float32_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float32_t"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_sub_f32(\n    const float32_t *input_1_vect,\n    const float32_t *input_2_vect,\n    float32_t *output,\n    float32_t out_activation_min,\n    float32_t out_activation_max,\n    int32_t block_size\n)",
              "source": {
                "line": 746,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L746"
              },
              "summary": "Elementwise subtract with optional output clamp."
            },
            {
              "description": "Elementwise absolute value.",
              "examples": [],
              "id": "arm_nn_abs_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_abs_f32",
              "params": [
                {
                  "description": "Pointer to the input vector.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_abs_f32(const float32_t *input, float32_t *output, int32_t block_size)",
              "source": {
                "line": 762,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L762"
              },
              "summary": "Elementwise absolute value."
            },
            {
              "description": "Fill a float32 vector with one value.\n\nBit copy of `value` into every element (vector splat / plain stores), so a NaN fill value lands bit-exact, sign and payload included. Not named arm_fill_f32: CMSIS-DSP exports that symbol.",
              "examples": [],
              "id": "arm_nn_fill_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_fill_f32",
              "params": [
                {
                  "description": "Fill value.",
                  "direction": "in",
                  "name": "value",
                  "type": "float32_t"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of elements to write (0 is a no-op).",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_fill_f32(float32_t value, float32_t *output, int32_t block_size)",
              "source": {
                "line": 778,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L778"
              },
              "summary": "Fill a float32 vector with one value."
            },
            {
              "description": "Elementwise multiply with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
              "examples": [],
              "id": "arm_elementwise_mul_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_mul_f32",
              "params": [
                {
                  "description": "Pointer to the first input vector.",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the second input vector.",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float32_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float32_t"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_mul_f32(\n    const float32_t *input_1_vect,\n    const float32_t *input_2_vect,\n    float32_t *output,\n    float32_t out_activation_min,\n    float32_t out_activation_max,\n    int32_t block_size\n)",
              "source": {
                "line": 805,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L805"
              },
              "summary": "Elementwise multiply with optional output clamp."
            },
            {
              "description": "Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.\n\nThe result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.",
              "examples": [],
              "id": "arm_minimum_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_minimum_f32",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Input 1, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input 2, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_minimum_f32(\n    const cmsis_nn_context *ctx,\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 858,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L858"
              },
              "summary": "Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape."
            },
            {
              "description": "Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.\n\nThe result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.",
              "examples": [],
              "id": "arm_maximum_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_maximum_f32",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Input 1, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input 2, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_maximum_f32(\n    const cmsis_nn_context *ctx,\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 894,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L894"
              },
              "summary": "Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape."
            },
            {
              "description": "Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and `output_dims` must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.\n\nNumerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged  a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as `scalar - element`, not as the negation of `element - scalar`, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.",
              "examples": [],
              "id": "arm_elementwise_sub_broadcast_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_sub_broadcast_f32",
              "params": [
                {
                  "description": "Minuend, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Subtrahend, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float32_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_sub_broadcast_f32(\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t out_activation_min,\n    float32_t out_activation_max\n)",
              "source": {
                "line": 933,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L933"
              },
              "summary": "Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp."
            },
            {
              "description": "Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.",
              "examples": [],
              "id": "arm_elementwise_add_broadcast_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_add_broadcast_f32",
              "params": [
                {
                  "description": "First input, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Second input, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float32_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "arm_cmsis_nn_status arm_elementwise_add_broadcast_f32(\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t out_activation_min,\n    float32_t out_activation_max\n)",
              "source": {
                "line": 957,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L957"
              },
              "summary": "Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp."
            },
            {
              "description": "Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.",
              "examples": [],
              "id": "arm_elementwise_mul_broadcast_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_mul_broadcast_f32",
              "params": [
                {
                  "description": "First input, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Second input, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float32_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "arm_cmsis_nn_status arm_elementwise_mul_broadcast_f32(\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t out_activation_min,\n    float32_t out_activation_max\n)",
              "source": {
                "line": 981,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L981"
              },
              "summary": "Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp."
            },
            {
              "description": "Elementwise square root.\n\nThe value path is scalar on every toolchain, because Helium has no vector square root. armclang and ATfE do vectorize the surrounding special-value classification; the results are bit-identical to the GCC scalar build, verified by executing both toolchains' objects (#295). Normal positive inputs evaluate `sqrtf(x)`, which IEEE 754 makes correctly rounded, so results are bit-exact to a float64 reference. Subnormal inputs follow FPSCR.FZ: where flush-to-zero is set - the Corstone-300 FVP default, and any host binary linked at -Ofast, where crtfastmath sets DAZ and FTZ - a subnormal input reads as zero and the result is +0. The float16 pair is immune, because it widens to a normal float32 first. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).",
              "examples": [],
              "id": "arm_nn_sqrt_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_sqrt_f32",
              "params": [
                {
                  "description": "Pointer to the input vector.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output vector; may alias `input`.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_sqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)",
              "source": {
                "line": 1012,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1012"
              },
              "summary": "Elementwise square root."
            },
            {
              "description": "Elementwise reciprocal square root, `1 / sqrt(x)`.\n\nSame value path as arm_nn_sqrt_f32, including its FPSCR.FZ behaviour on subnormal inputs, where the result is +Inf. Normal positive inputs evaluate `1.0f / sqrtf(x)` in float32: two IEEE roundings, so the result is within 1 ulp of the correctly rounded value (measured against a float64 reference; `x = 4^k` is exact). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7FC00000, NaN -> the same NaN with the quiet bit set (sign and payload kept).",
              "examples": [],
              "id": "arm_rsqrt_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_rsqrt_f32",
              "params": [
                {
                  "description": "Pointer to the input vector.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output vector; may alias `input`.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_rsqrt_f32(const float32_t *input, float32_t *output, int32_t block_size)",
              "source": {
                "line": 1031,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1031"
              },
              "summary": "Elementwise reciprocal square root, 1 / sqrt(x)."
            },
            {
              "description": "Elementwise add with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf + (-Inf) for add), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
              "examples": [],
              "id": "arm_elementwise_add_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_add_f16",
              "params": [
                {
                  "description": "Pointer to the first input vector.",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the second input vector.",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float16_t"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_add_f16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    float16_t out_activation_min,\n    float16_t out_activation_max,\n    int32_t block_size\n)",
              "source": {
                "line": 2815,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2815"
              },
              "summary": "Elementwise add with optional output clamp."
            },
            {
              "description": "Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate `arm_elementwise_add_f16()`. New code should call `arm_elementwise_add_f16()` instead.\n\nThis entry does NOT share the contract of `arm_elementwise_add_f16()`:\n\n:::caution\nNo argument validation is performed. A NULL `input_1_vect`, `input_2_vect` or `output` is dereferenced rather than reported. A `block_size` of 0 writes nothing and still returns `ARM_CMSIS_NN_SUCCESS`, where `arm_elementwise_add_f16()` returns `ARM_CMSIS_NN_ARG_ERROR`.\n\n:::\n\n:::note\nThe clamp does not propagate NaN. Both the Helium path (`vminnm`/`vmaxnm`) and the scalar path (the non-propagating `MIN`/`MAX` clamp helper) bound against `out_activation_max` first, so a NaN produced by the addition comes back as `out_activation_max`. `arm_elementwise_add_f16()` documents TensorFlow Lite NaN propagation; this entry does not implement it.\n\n:::",
              "examples": [],
              "id": "arm_elementwise_add_fp16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_add_fp16",
              "params": [
                {
                  "description": "Pointer to the first input vector. Must not be NULL.",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the second input vector. Must not be NULL.",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output vector. Must not be NULL.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const float16_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const float16_t"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` unconditionally."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_add_fp16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    const float16_t out_activation_min,\n    const float16_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 2846,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2846"
              },
              "summary": "Legacy float16 elementwise add with fused clamp, kept only for source compatibility with callers that predate armelementwiseaddf16()."
            },
            {
              "description": "Elementwise subtract with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as Inf - Inf for subtract), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
              "examples": [],
              "id": "arm_elementwise_sub_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_sub_f16",
              "params": [
                {
                  "description": "Pointer to the first input vector (minuend).",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the second input vector (subtrahend).",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float16_t"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_sub_f16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    float16_t out_activation_min,\n    float16_t out_activation_max,\n    int32_t block_size\n)",
              "source": {
                "line": 2856,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2856"
              },
              "summary": "Elementwise subtract with optional output clamp."
            },
            {
              "description": "Elementwise squared difference of two float16 vectors.\n\nEach output element is calculated as `(input_1_vect[i] - input_2_vect[i])^2`.",
              "examples": [],
              "id": "arm_elementwise_squared_difference_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_squared_difference_f16",
              "params": [
                {
                  "description": "Pointer to the first input vector.",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the second input vector.",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when an input/output pointer is NULL or `block_size` is less than 1."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_squared_difference_f16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    int32_t block_size\n)",
              "source": {
                "line": 2876,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2876"
              },
              "summary": "Elementwise squared difference of two float16 vectors."
            },
            {
              "description": "Elementwise absolute value.",
              "examples": [],
              "id": "arm_nn_abs_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_abs_f16",
              "params": [
                {
                  "description": "Pointer to the input vector.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_abs_f16(const float16_t *input, float16_t *output, int32_t block_size)",
              "source": {
                "line": 2884,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2884"
              },
              "summary": "Elementwise absolute value."
            },
            {
              "description": "Fill a float16 vector with one value; bit copy of `value`, NaN payload included.",
              "examples": [],
              "id": "arm_nn_fill_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_fill_f16",
              "params": [
                {
                  "description": "Fill value.",
                  "direction": "in",
                  "name": "value",
                  "type": "float16_t"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of elements to write (0 is a no-op).",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or `output` is NULL with a non-zero `block_size`."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_fill_f16(float16_t value, float16_t *output, int32_t block_size)",
              "source": {
                "line": 2896,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2896"
              },
              "summary": "Fill a float16 vector with one value; bit copy of value, NaN payload included."
            },
            {
              "description": "Split a float32 tensor of any rank into several tensors along one axis.\n\nInverse of arm_concatenation_f32; per-split lengths also cover SPLIT_V. Output `s` has the input shape with `input_shape`[axis] replaced by `split_dims`[s]. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input. A dimension of 0 is accepted and copies nothing.",
              "examples": [],
              "id": "arm_split_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_split_f16",
              "params": [
                {
                  "description": "Pointer to the flattened (row-major) input.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of dimensions in `input_shape` (>= 1).",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const int32_t"
                },
                {
                  "description": "Input shape; `input_shape`[axis] must equal the sum of `split_dims`.",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Axis to split along (0 <= axis < input_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of outputs (>= 1).",
                  "direction": "in",
                  "name": "num_splits",
                  "type": "const int32_t"
                },
                {
                  "description": "Array of length `num_splits:` each output's extent along `axis`.",
                  "direction": "in",
                  "name": "split_dims",
                  "type": "const int32_t *"
                },
                {
                  "description": "Array of `num_splits` pointers to the flattened outputs.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *const *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, split entry, split sum, NULL pointer or an element count above INT32_MAX."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_split_f16(\n    const float16_t *input_data,\n    const int32_t input_dims,\n    const int32_t *input_shape,\n    const int32_t axis,\n    const int32_t num_splits,\n    const int32_t *split_dims,\n    float16_t *const *output_data\n)",
              "source": {
                "line": 2923,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2923"
              },
              "summary": "Split a float32 tensor of any rank into several tensors along one axis."
            },
            {
              "description": "Strided slice for float32 data (pure copy, TensorFlow Lite compatible).",
              "examples": [],
              "id": "arm_strided_slice_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_strided_slice_f16",
              "params": [
                {
                  "description": "Pointer to input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to output tensor.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Begin dimensions for slicing.",
                  "direction": "in",
                  "name": "begin_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Stride dimensions for slicing.",
                  "direction": "in",
                  "name": "stride_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *const"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_strided_slice_f16(\n    const float16_t *input_data,\n    float16_t *output_data,\n    const cmsis_nn_dims *const input_dims,\n    const cmsis_nn_dims *const begin_dims,\n    const cmsis_nn_dims *const stride_dims,\n    const cmsis_nn_dims *const output_dims\n)",
              "source": {
                "line": 2934,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2934"
              },
              "summary": "Strided slice for float32 data (pure copy, TensorFlow Lite compatible)."
            },
            {
              "description": "Elementwise multiply with optional output clamp.\n\nNaN propagates through the clamp (TensorFlow Lite semantics): a quiet NaN in either input operand, or a NaN produced by the arithmetic itself (such as 0 * Inf for multiply), yields a NaN at that output element. This holds at every optimization level, on the toolchains this project gates (see the Testing & Verification guide, docs/guides/verification.md), including the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt): the clamp classifies NaN on the integer bit pattern of the value rather than with a floating-point compare, and the -ffinite-math-only that -Ofast implies grants no license to fold integer arithmetic. Verified by host execution and by disassembly on gated Arm GNU Toolchain 14.3.Rel1 (where the unguarded form demonstrably folds) and 13.x/15.x and armclang 6.23; ATfE unexamined  cross-toolchain on-target execution is #340's scope. See issues #333 and #334. Only the NaN-ness of the element is guaranteed, not a particular NaN payload. Infinities that are not NaN still clamp to the activation bounds (and pass through unchanged under the +/-INFINITY \"no clamp\" idiom).",
              "examples": [],
              "id": "arm_elementwise_mul_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_mul_f16",
              "params": [
                {
                  "description": "Pointer to the first input vector.",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the second input vector.",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float16_t"
                },
                {
                  "description": "Number of elements to process.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_mul_f16(\n    const float16_t *input_1_vect,\n    const float16_t *input_2_vect,\n    float16_t *output,\n    float16_t out_activation_min,\n    float16_t out_activation_max,\n    int32_t block_size\n)",
              "source": {
                "line": 2944,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2944"
              },
              "summary": "Elementwise multiply with optional output clamp."
            },
            {
              "description": "Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.\n\nThe result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.",
              "examples": [],
              "id": "arm_minimum_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_minimum_f16",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Input 1, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input 2, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_minimum_f16(\n    const cmsis_nn_context *ctx,\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2954,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2954"
              },
              "summary": "Elementwise minimum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape."
            },
            {
              "description": "Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and `output_dims` must be their broadcast shape.\n\nThe result for a tie between zeros of opposite sign, and for any non-finite input, is unspecified. The Helium leg is VMAXNM / VMINNM, which implement IEEE maxNum / minNum: they break a zero tie by sign - maximum returns +0.0, minimum returns -0.0 - and they suppress NaN, returning the non-NaN operand and a default quiet NaN when both operands are NaN. The scalar leg breaks the tie by operand position instead, and which position wins is not fixed either: the shipped -Ofast (CMSIS_OPTIMIZATION_LEVEL in the top-level CMakeLists.txt) implies -fno-signed-zeros and -ffinite-math-only, which license the compiler to answer a zero tie or a NaN either way, and the answer measurably differs between build targets, between optimization levels, and between the contiguous and the broadcast-scalar loop of the same build. Both zero answers compare equal to zero, so the difference is invisible to anything that is not bit-exact; a caller that cares about the sign of a zero, or about NaN, must screen its inputs rather than rely on either leg. See issue #316, and #333 for the same -ffinite-math-only caveat on the elementwise family.",
              "examples": [],
              "id": "arm_maximum_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_maximum_f16",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Input 1, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input 2, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_maximum_f16(\n    const cmsis_nn_context *ctx,\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2965,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2965"
              },
              "summary": "Elementwise maximum with TensorFlow Lite NHWC broadcasting: each dimension of the two inputs must be equal or 1, and outputdims must be their broadcast shape."
            },
            {
              "description": "Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcasting follows the NumPy / TensorFlow Lite rule per dimension: each of n, h, w and c of the two inputs must be equal or 1, a dimension of 1 is repeated along that axis, and `output_dims` must be the elementwise maximum of the two input shapes. A dimension of 0 or less is rejected.\n\nNumerics are those of arm_elementwise_sub_f32 applied to the materialised broadcast operands: identical arithmetic and clamp on every path, so on the shipped Cortex-M legs (M4, M55) the output is bit-identical to that kernel, NaN payload aside, and its NaN contract holds here unchanged  a NaN in either operand, or one produced by the arithmetic, propagates through the clamp at every optimization level, while non-NaN infinities clamp to the bounds. On other hosts built with -fno-signed-zeros the sign of a zero that ties with a zero clamp bound is compiler-licensed and may differ between this walk and the flat loop. The bounds must be ordered and non-NaN. When input 1 is the broadcast scalar the result is computed as `scalar - element`, not as the negation of `element - scalar`, which differs at a zero result; whether the sign of a zero survives is then subject to the same -fno-signed-zeros license the shipped -Ofast grants the compiler on the flat kernels.\n\nHalf-precision twin: the numerics are those of arm_elementwise_sub_f16 on the materialised operands.",
              "examples": [],
              "id": "arm_elementwise_sub_broadcast_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_sub_broadcast_f16",
              "params": [
                {
                  "description": "Minuend, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Subtrahend, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL, a dimension is not positive, the shapes are not broadcast-compatible, or the output shape is not their broadcast shape. Nothing is written on error."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_sub_broadcast_f16(\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t out_activation_min,\n    float16_t out_activation_max\n)",
              "source": {
                "line": 2978,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2978"
              },
              "summary": "Elementwise subtract with TensorFlow Lite NHWC broadcasting and an output clamp."
            },
            {
              "description": "Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_add_f32 on the materialised broadcast operands, including its NaN contract.\n\nHalf-precision twin: the numerics are those of arm_elementwise_add_f16 on the materialised operands.",
              "examples": [],
              "id": "arm_elementwise_add_broadcast_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_add_broadcast_f16",
              "params": [
                {
                  "description": "First input, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Second input, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "arm_cmsis_nn_status arm_elementwise_add_broadcast_f16(\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t out_activation_min,\n    float16_t out_activation_max\n)",
              "source": {
                "line": 2992,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2992"
              },
              "summary": "Elementwise add with TensorFlow Lite NHWC broadcasting and an output clamp."
            },
            {
              "description": "Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp.\n\nBroadcast rules, argument checking and return values as for arm_elementwise_sub_broadcast_f32; numerics are those of arm_elementwise_mul_f32 on the materialised broadcast operands, including its NaN contract.\n\nHalf-precision twin: the numerics are those of arm_elementwise_mul_f16 on the materialised operands.",
              "examples": [],
              "id": "arm_elementwise_mul_broadcast_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_mul_broadcast_f16",
              "params": [
                {
                  "description": "First input, NHWC, sized by `input_1_dims`.",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 1.",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Second input, NHWC, sized by `input_2_dims`.",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of input 2.",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output, NHWC, sized by `output_dims`.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Broadcast output dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Minimum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Maximum output clamp value.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "arm_cmsis_nn_status arm_elementwise_mul_broadcast_f16(\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t out_activation_min,\n    float16_t out_activation_max\n)",
              "source": {
                "line": 3006,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3006"
              },
              "summary": "Elementwise multiply with TensorFlow Lite NHWC broadcasting and an output clamp."
            },
            {
              "description": "Elementwise square root of a float16 tensor.\n\nThe value path is scalar on every toolchain, because Helium has no vector square root; armclang and ATfE vectorize the surrounding classification into an MVE loop and produce bit-identical results, verified by executing their objects (#295). Each element is widened to float32, `sqrtf` is evaluated there and the result is rounded once to float16. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 square root (0 ulp, #295). Widening first also makes this pair immune to FPSCR.FZ, which flushes float32 subnormals in the f32 pair. Special values are decided on the bit pattern and returned as literals, independent of -ffinite-math-only and FPSCR.DN: +0 -> +0, -0 -> -0, +Inf -> +Inf, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).",
              "examples": [],
              "id": "arm_nn_sqrt_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_sqrt_f16",
              "params": [
                {
                  "description": "Pointer to the input tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output tensor; may alias `input`.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of tensor elements.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_sqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)",
              "source": {
                "line": 3036,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3036"
              },
              "summary": "Elementwise square root of a float16 tensor."
            },
            {
              "description": "Elementwise reciprocal square root of a float16 tensor, `1 / sqrt(x)`.\n\nSame value path as arm_nn_sqrt_f16: widen to float32, evaluate `1.0f / sqrtf(x)` there, round once to float16, and so also immune to FPSCR.FZ. Verified exhaustively: for every positive finite float16 input, subnormals included, the result is the correctly rounded float16 of the float64 reciprocal square root (0 ulp, #295). Special values are decided on the bit pattern and returned as literals: +0 -> +Inf, -0 -> -Inf, +Inf -> +0, negative (including -Inf) -> quiet NaN 0x7E00, NaN -> the same NaN with the quiet bit set (sign and payload kept).",
              "examples": [],
              "id": "arm_rsqrt_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_rsqrt_f16",
              "params": [
                {
                  "description": "Pointer to the input tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output tensor; may alias `input`.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of tensor elements.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_rsqrt_f16(const float16_t *input, float16_t *output, int32_t block_size)",
              "source": {
                "line": 3055,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3055"
              },
              "summary": "Elementwise reciprocal square root of a float16 tensor, 1 / sqrt(x)."
            }
          ]
        },
        {
          "description": "Internal Support functions. Not intended to be called direclty by a CMSIS-NN user.",
          "name": "Private",
          "path": "heliaCORE.groupSupport",
          "submodules": [],
          "summary": "Internal Support functions.",
          "symbols": [
            {
              "description": "Polynomial coefficients used by the float32 MVE exp approximation.",
              "examples": [],
              "id": "arm_nn_exp_poly_coeffs_f32",
              "kind": "attribute",
              "language": "c",
              "members": [],
              "name": "arm_nn_exp_poly_coeffs_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "const float32_t arm_nn_exp_poly_coeffs_f32[8]",
              "source": {
                "line": 70,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L70"
              },
              "summary": "Polynomial coefficients used by the float32 MVE exp approximation."
            },
            {
              "description": "LUT for `2^(i/256)` used by the float32 LUT softmax approximation.\n\nStores 257 samples for `i = 0..256` so interpolation can safely read `lut[idx + 1]` while indexing the 256 fractional segments.",
              "examples": [],
              "id": "arm_nn_exp2_lut_f32",
              "kind": "attribute",
              "language": "c",
              "members": [],
              "name": "arm_nn_exp2_lut_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "const float32_t arm_nn_exp2_lut_f32[257]",
              "source": {
                "line": 78,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L78"
              },
              "summary": "LUT for 2^(i/256) used by the float32 LUT softmax approximation."
            },
            {
              "description": "Floor of `x` as an int32_t.\n\nPrecondition: `x` must already be reduced to the int32_t range and must not be NaN  the float-to-int conversion below is undefined otherwise. The only caller, arm_nn_softmax_exp_lut_f32(), guarantees this by clamping its input to [-80, 80] (NaN included, see there) before scaling by log2(e), which bounds `x` to +/-116.",
              "examples": [],
              "id": "arm_nn_softmax_floor_to_int_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_floor_to_int_f32",
              "params": [
                {
                  "description": "Value to floor.",
                  "direction": "in",
                  "name": "x",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Largest int32_t not greater than `x`."
                }
              ],
              "signature": "static int32_t arm_nn_softmax_floor_to_int_f32(float32_t x)",
              "source": {
                "line": 92,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L92"
              },
              "summary": "Floor of x as an int32t."
            },
            {
              "description": "Reinterpret a 32-bit pattern as a float32.",
              "examples": [],
              "id": "arm_nn_softmax_fp32_from_bits",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_fp32_from_bits",
              "params": [
                {
                  "description": "IEEE-754 binary32 bit pattern.",
                  "direction": "in",
                  "name": "bits",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The float32 value with the bit pattern `bits`."
                }
              ],
              "signature": "static float32_t arm_nn_softmax_fp32_from_bits(uint32_t bits)",
              "source": {
                "line": 104,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L104"
              },
              "summary": "Reinterpret a 32-bit pattern as a float32."
            },
            {
              "description": "Compute `2^n` as a float32 by building the exponent field directly.",
              "examples": [],
              "id": "arm_nn_softmax_exp2i_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_exp2i_f32",
              "params": [
                {
                  "description": "Integer exponent. Clamped to the normal float32 exponent range `[-126, 127]`.",
                  "direction": "in",
                  "name": "n",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`2^n` as a float32."
                }
              ],
              "signature": "static float32_t arm_nn_softmax_exp2i_f32(int32_t n)",
              "source": {
                "line": 121,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L121"
              },
              "summary": "Compute 2^n as a float32 by building the exponent field directly."
            },
            {
              "description": "Taylor/Estrin exp approximation for float32 softmax helpers.\n\nThe polynomial is evaluated on r in [-ln2/2, ln2/2]. Coefficients come from the Maclaurin series of exp(r): exp(r) ~= 1 + r + r^2/2! + r^3/3! + r^4/4! + r^5/5! + r^6/6! Grouped via Estrin to reduce dependency depth: p = (1 + r) + r^2*(1/2 + r/6) + r^4*(1/24 + r/120) + r^6*(1/720)\n\nRange reduction follows: x = n * ln(2) + r, exp(x) = exp(r) * 2^n",
              "examples": [],
              "id": "arm_nn_softmax_exp_taylor_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_exp_taylor_f32",
              "params": [
                {
                  "description": "Exponent argument. Clamped to `[-80, 80]` before evaluation.",
                  "direction": "in",
                  "name": "x",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Approximation of `exp(x)`."
                }
              ],
              "signature": "static float32_t arm_nn_softmax_exp_taylor_f32(float32_t x)",
              "source": {
                "line": 147,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L147"
              },
              "summary": "Taylor/Estrin exp approximation for float32 softmax helpers."
            },
            {
              "description": "LUT-based exp approximation for float32 softmax helpers.\n\nSplits `x * log2(e)` into an integer part handled by arm_nn_softmax_exp2i_f32() and a fractional part interpolated linearly from `arm_nn_exp2_lut_f32`.",
              "examples": [],
              "id": "arm_nn_softmax_exp_lut_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_exp_lut_f32",
              "params": [
                {
                  "description": "Exponent argument. Clamped to `[-80, 80]` before evaluation; NaN is flushed to `80`.",
                  "direction": "in",
                  "name": "x",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Approximation of `exp(x)`."
                }
              ],
              "signature": "static float32_t arm_nn_softmax_exp_lut_f32(float32_t x)",
              "source": {
                "line": 184,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L184"
              },
              "summary": "LUT-based exp approximation for float32 softmax helpers."
            },
            {
              "description": "Scalar exp approximation used by the float32 softmax paths.\n\nDispatches to arm_nn_softmax_exp_taylor_f32() when `ARM_NN_USE_EXP_TAYLOR` is defined and to arm_nn_softmax_exp_lut_f32() otherwise.",
              "examples": [],
              "id": "arm_nn_softmax_exp_scalar_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_exp_scalar_f32",
              "params": [
                {
                  "description": "Exponent argument.",
                  "direction": "in",
                  "name": "x",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Approximation of `exp(x)`."
                }
              ],
              "signature": "static float32_t arm_nn_softmax_exp_scalar_f32(float32_t x)",
              "source": {
                "line": 254,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L254"
              },
              "summary": "Scalar exp approximation used by the float32 softmax paths."
            },
            {
              "description": "LUT for tanh(x) sampled over `x in [0, 6]` for float32 helpers.\n\nStores 385 samples so interpolation can safely read `lut[idx + 1]` while indexing the 384 fractional segments across the interval. The grid spacing (`6/384 == 1/64`) matches the earlier 257-entry `[0, 4]` table, so entries `0..256` are bit-identical to it and the index multiplier is unchanged. Generated by `scripts/gen_tanh_lut_f32.py`.",
              "examples": [],
              "id": "arm_nn_tanh_lut_f32",
              "kind": "attribute",
              "language": "c",
              "members": [],
              "name": "arm_nn_tanh_lut_f32",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "const float32_t arm_nn_tanh_lut_f32[385]",
              "source": {
                "line": 276,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L276"
              },
              "summary": "LUT for tanh(x) sampled over x in [0, 6] for float32 helpers."
            },
            {
              "description": "Copy a float32 vector.",
              "examples": [],
              "id": "arm_memcpy_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memcpy_f32",
              "params": [
                {
                  "description": "Destination buffer.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float32_t *"
                },
                {
                  "description": "Source buffer.",
                  "direction": "in",
                  "name": "src",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of elements to copy.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memcpy_f32(float32_t *dst, const float32_t *src, uint32_t block_size)",
              "source": {
                "line": 311,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L311"
              },
              "summary": "Copy a float32 vector."
            },
            {
              "description": "Set a float32 vector to a constant value.",
              "examples": [],
              "id": "arm_memset_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memset_f32",
              "params": [
                {
                  "description": "Destination buffer.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float32_t *"
                },
                {
                  "description": "Fill value.",
                  "direction": "in",
                  "name": "val",
                  "type": "const float32_t"
                },
                {
                  "description": "Number of elements to write.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memset_f32(float32_t *dst, const float32_t val, uint32_t block_size)",
              "source": {
                "line": 334,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L334"
              },
              "summary": "Set a float32 vector to a constant value."
            },
            {
              "description": "Specialized NHWC depthwise 1D kernel for `k=3`, `ch_mult=1` (float32).",
              "examples": [],
              "id": "arm_nn_depthwise_conv1d_k3_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv1d_k3_nhwc_f32",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of input (and output) channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Depthwise weights with shape `[3][in_c]`.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float32_t *"
                },
                {
                  "description": "Optional bias vector of `in_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][in_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float32_t *"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_depthwise_conv1d_k3_nhwc_f32(\n    const float32_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float32_t *kernel,\n    const float32_t *b,\n    float32_t *out,\n    int32_t out_w\n)",
              "source": {
                "line": 367,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L367"
              },
              "summary": "Specialized NHWC depthwise 1D kernel for k=3, chmult=1 (float32)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=5` (float32).",
              "examples": [],
              "id": "arm_nn_conv1d_k5_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k5_nhwc_f32",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights with shape `[out_c][5][in_c]`.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float32_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+4`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k5_nhwc_f32(\n    const float32_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float32_t *kernel,\n    const float32_t *b,\n    float32_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 387,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L387"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=5 (float32)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=5` (float32, packed weights).\n\nThe packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f32`, i.e. `[(5 * in_c)][out_c_block_of_4]`.",
              "examples": [],
              "id": "arm_nn_conv1d_k5_packed_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k5_packed_f32",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights packed in output-channel blocks of 4 as described above.",
                  "direction": "in",
                  "name": "kernel_packed",
                  "type": "const float32_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+4`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k5_packed_f32(\n    const float32_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float32_t *kernel_packed,\n    const float32_t *b,\n    float32_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 411,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L411"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=5 (float32, packed weights)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=3` (float32).",
              "examples": [],
              "id": "arm_nn_conv1d_k3_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k3_nhwc_f32",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights with shape `[out_c][3][in_c]`.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float32_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k3_nhwc_f32(\n    const float32_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float32_t *kernel,\n    const float32_t *b,\n    float32_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 432,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L432"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=3 (float32)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=3` (float32, packed weights).\n\nThe packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f32`, i.e. `[(3 * in_c)][out_c_block_of_4]`.",
              "examples": [],
              "id": "arm_nn_conv1d_k3_packed_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k3_packed_f32",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights packed in output-channel blocks of 4 as described above.",
                  "direction": "in",
                  "name": "kernel_packed",
                  "type": "const float32_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k3_packed_f32(\n    const float32_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float32_t *kernel_packed,\n    const float32_t *b,\n    float32_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 456,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L456"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=3 (float32, packed weights)."
            },
            {
              "description": "Specialized NHWC max-pool 1D kernel for `k=3`, `s=3` (float32).",
              "examples": [],
              "id": "arm_nn_maxpool1d_k3s3_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_maxpool1d_k3s3_nhwc_f32",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][in_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float32_t *"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `3*ow..3*ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_maxpool1d_k3s3_nhwc_f32(const float32_t *x_nhwc, int32_t in_c, int32_t in_w, float32_t *out, int32_t out_w)",
              "source": {
                "line": 474,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L474"
              },
              "summary": "Specialized NHWC max-pool 1D kernel for k=3, s=3 (float32)."
            },
            {
              "description": "Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` without output clamp (float32).",
              "examples": [],
              "id": "arm_nn_maxpool1d_k2s2_nhwc_noclip_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_maxpool1d_k2s2_nhwc_noclip_f32",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][in_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float32_t *"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `2*ow..2*ow+1`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_maxpool1d_k2s2_nhwc_noclip_f32(\n    const float32_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    float32_t *out,\n    int32_t out_w\n)",
              "source": {
                "line": 489,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L489"
              },
              "summary": "Specialized NHWC max-pool 1D kernel for k=2, s=2 without output clamp (float32)."
            },
            {
              "description": "Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` with clamp (float32).",
              "examples": [],
              "id": "arm_nn_maxpool1d_k2s2_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_maxpool1d_k2s2_nhwc_f32",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][in_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float32_t *"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `2*ow..2*ow+1`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                },
                {
                  "description": "Lower clamp bound applied to `out`.",
                  "direction": "in",
                  "name": "act_min",
                  "type": "float32_t"
                },
                {
                  "description": "Upper clamp bound applied to `out`.",
                  "direction": "in",
                  "name": "act_max",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_maxpool1d_k2s2_nhwc_f32(\n    const float32_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    float32_t *out,\n    int32_t out_w,\n    float32_t act_min,\n    float32_t act_max\n)",
              "source": {
                "line": 506,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L506"
              },
              "summary": "Specialized NHWC max-pool 1D kernel for k=2, s=2 with clamp (float32)."
            },
            {
              "description": "Matrix multiply with non-transposed lhs and transposed rhs rows (float32).",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_t_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_t_f32",
              "params": [
                {
                  "description": "Left-hand matrix stored row-major.",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const float32_t *"
                },
                {
                  "description": "Right-hand matrix stored row-major, one row per output channel.",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const float32_t *"
                },
                {
                  "description": "Optional bias vector.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output matrix.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of rows in `lhs`.",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of rows in `rhs`.",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of columns in `rhs`.",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "int32_t"
                },
                {
                  "description": "Output row stride, expressed in elements.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "int32_t"
                },
                {
                  "description": "Lower clamp bound.",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "float32_t"
                },
                {
                  "description": "Upper clamp bound.",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f32(\n    const float32_t *lhs,\n    const float32_t *rhs,\n    const float32_t *bias,\n    float32_t *dst,\n    int32_t lhs_rows,\n    int32_t rhs_rows,\n    int32_t rhs_cols,\n    int32_t row_address_offset,\n    float32_t activation_min,\n    float32_t activation_max\n)",
              "source": {
                "line": 529,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L529"
              },
              "summary": "Matrix multiply with non-transposed lhs and transposed rhs rows (float32)."
            },
            {
              "description": "Matrix multiply with non-transposed lhs and packed non-transposed rhs (float32).",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_n_packed_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_n_packed_f32",
              "params": [
                {
                  "description": "Left-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`.",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const float32_t *"
                },
                {
                  "description": "Right-hand matrix with logical shape `[rhs_cols, rhs_rows]`, packed in column blocks of 4. The final block uses the same packed stride and inactive tail lanes are ignored.",
                  "direction": "in",
                  "name": "rhs_packed",
                  "type": "const float32_t *"
                },
                {
                  "description": "Optional bias vector.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output matrix.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of rows in `lhs`.",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of logical output columns in the unpacked rhs matrix.",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Shared reduction dimension `K`.",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "int32_t"
                },
                {
                  "description": "Output row stride, expressed in elements.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "int32_t"
                },
                {
                  "description": "Lower clamp bound.",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "float32_t"
                },
                {
                  "description": "Upper clamp bound.",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "float32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f32(\n    const float32_t *lhs,\n    const float32_t *rhs_packed,\n    const float32_t *bias,\n    float32_t *dst,\n    int32_t lhs_rows,\n    int32_t rhs_rows,\n    int32_t rhs_cols,\n    int32_t row_address_offset,\n    float32_t activation_min,\n    float32_t activation_max\n)",
              "source": {
                "line": 556,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L556"
              },
              "summary": "Matrix multiply with non-transposed lhs and packed non-transposed rhs (float32)."
            },
            {
              "description": "Pack a single convolution patch into one row of a contiguous float32 patch matrix.\n\nDevelopers familiar with im2row/im2col terminology can think of this as packing one output patch into one row.",
              "examples": [],
              "id": "arm_nn_pack_conv_patch_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_pack_conv_patch_f32",
              "params": [
                {
                  "description": "Input tensor for one batch in NHWC layout with shape `[in_h][in_w][in_c]`.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Input height.",
                  "direction": "in",
                  "name": "in_h",
                  "type": "int32_t"
                },
                {
                  "description": "Input width.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Kernel height.",
                  "direction": "in",
                  "name": "kernel_h",
                  "type": "int32_t"
                },
                {
                  "description": "Kernel width.",
                  "direction": "in",
                  "name": "kernel_w",
                  "type": "int32_t"
                },
                {
                  "description": "Vertical stride.",
                  "direction": "in",
                  "name": "stride_h",
                  "type": "int32_t"
                },
                {
                  "description": "Horizontal stride.",
                  "direction": "in",
                  "name": "stride_w",
                  "type": "int32_t"
                },
                {
                  "description": "Top padding.",
                  "direction": "in",
                  "name": "pad_h",
                  "type": "int32_t"
                },
                {
                  "description": "Left padding.",
                  "direction": "in",
                  "name": "pad_w",
                  "type": "int32_t"
                },
                {
                  "description": "Vertical dilation.",
                  "direction": "in",
                  "name": "dilation_h",
                  "type": "int32_t"
                },
                {
                  "description": "Horizontal dilation.",
                  "direction": "in",
                  "name": "dilation_w",
                  "type": "int32_t"
                },
                {
                  "description": "Output row index of the patch to pack.",
                  "direction": "in",
                  "name": "out_y",
                  "type": "int32_t"
                },
                {
                  "description": "Output column index of the patch to pack.",
                  "direction": "in",
                  "name": "out_x",
                  "type": "int32_t"
                },
                {
                  "description": "Value written for taps that fall outside the input.",
                  "direction": "in",
                  "name": "pad_value",
                  "type": "float32_t"
                },
                {
                  "description": "Destination row of `kernel_h * kernel_w * in_c` elements, ordered `[kernel_h][kernel_w][in_c]`.",
                  "direction": "out",
                  "name": "patch_row",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_pack_conv_patch_f32(\n    const float32_t *input,\n    int32_t in_h,\n    int32_t in_w,\n    int32_t in_c,\n    int32_t kernel_h,\n    int32_t kernel_w,\n    int32_t stride_h,\n    int32_t stride_w,\n    int32_t pad_h,\n    int32_t pad_w,\n    int32_t dilation_h,\n    int32_t dilation_w,\n    int32_t out_y,\n    int32_t out_x,\n    float32_t pad_value,\n    float32_t *patch_row\n)",
              "source": {
                "line": 590,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L590"
              },
              "summary": "Pack a single convolution patch into one row of a contiguous float32 patch matrix."
            },
            {
              "description": "Specialized softmax helper for a single float32 row of length 2.",
              "examples": [],
              "id": "arm_nn_softmax_1x2_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_1x2_f32",
              "params": [
                {
                  "description": "Pointer to two contiguous float32 input values.",
                  "direction": "in",
                  "name": "in",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to two contiguous float32 output values.",
                  "direction": "out",
                  "name": "out",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_softmax_1x2_f32(const float32_t *in, float32_t *out)",
              "source": {
                "line": 613,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L613"
              },
              "summary": "Specialized softmax helper for a single float32 row of length 2."
            },
            {
              "description": "Blockwise float16 accumulation on the MVE legs (AmbiqAI/ns-cmsis-nn#586).\n\nA float16 accumulator lane sums at most ARM_NN_F16_ACC_BLOCK taps, in the kernel's tap order, before its partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once. The `_acc16` entries instantiate the same kernel bodies with ARM_NN_F16_ACC_BLOCK_NONE, which never folds.",
              "examples": [],
              "id": "ARM_NN_F16_ACC_BLOCK",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_F16_ACC_BLOCK",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_F16_ACC_BLOCK (32)",
              "source": {
                "line": 626,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L626"
              },
              "summary": "Blockwise float16 accumulation on the MVE legs (AmbiqAI/ns-cmsis-nn#586)."
            },
            {
              "description": "",
              "examples": [],
              "id": "ARM_NN_F16_ACC_BLOCK_NONE",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_F16_ACC_BLOCK_NONE",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_F16_ACC_BLOCK_NONE (INT32_MAX)",
              "source": {
                "line": 627,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L627"
              },
              "summary": ""
            },
            {
              "description": "Polynomial coefficients used by the float16 MVE exp approximation.\n\nThe float16 MVE helper evaluates the polynomial in widened float32 lanes, but it uses a dedicated coefficient table to keep the float16 path isolated from the float32 feature gate and softmax support stack.",
              "examples": [],
              "id": "arm_nn_exp_poly_coeffs_f16",
              "kind": "attribute",
              "language": "c",
              "members": [],
              "name": "arm_nn_exp_poly_coeffs_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "const float32_t arm_nn_exp_poly_coeffs_f16[8]",
              "source": {
                "line": 636,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L636"
              },
              "summary": "Polynomial coefficients used by the float16 MVE exp approximation."
            },
            {
              "description": "Quantized binary16 LUT for `2^(i/256)` used by float16 helpers.\n\nStores 257 samples for `i = 0..256` so interpolation can safely read `lut[idx + 1]` while indexing the 256 fractional segments.",
              "examples": [],
              "id": "arm_nn_exp2_lut_f16",
              "kind": "attribute",
              "language": "c",
              "members": [],
              "name": "arm_nn_exp2_lut_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "const uint16_t arm_nn_exp2_lut_f16[257]",
              "source": {
                "line": 644,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L644"
              },
              "summary": "Quantized binary16 LUT for 2^(i/256) used by float16 helpers."
            },
            {
              "description": "Quantized binary16 LUT for tanh(x) with `x in [0, 4]`.\n\nStores 257 samples so interpolation can safely read `lut[idx + 1]` while indexing the 256 fractional segments across the interval.",
              "examples": [],
              "id": "arm_nn_tanh_lut_f16",
              "kind": "attribute",
              "language": "c",
              "members": [],
              "name": "arm_nn_tanh_lut_f16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "const uint16_t arm_nn_tanh_lut_f16[257]",
              "source": {
                "line": 652,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L652"
              },
              "summary": "Quantized binary16 LUT for tanh(x) with x in [0, 4]."
            },
            {
              "description": "Reinterpret a 16-bit pattern as a float16.",
              "examples": [],
              "id": "arm_nn_softmax_fp16_from_bits",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_fp16_from_bits",
              "params": [
                {
                  "description": "IEEE-754 binary16 bit pattern.",
                  "direction": "in",
                  "name": "bits",
                  "type": "uint16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The float16 value with the bit pattern `bits`."
                }
              ],
              "signature": "static float16_t arm_nn_softmax_fp16_from_bits(uint16_t bits)",
              "source": {
                "line": 660,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L660"
              },
              "summary": "Reinterpret a 16-bit pattern as a float16."
            },
            {
              "description": "Floor of `x` as an int32_t.",
              "examples": [],
              "id": "arm_nn_softmax_floor_to_int_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_floor_to_int_f16",
              "params": [
                {
                  "description": "Value to floor. Must be finite and within the int32_t range.",
                  "direction": "in",
                  "name": "x",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Largest int32_t not greater than `x`."
                }
              ],
              "signature": "static int32_t arm_nn_softmax_floor_to_int_f16(float16_t x)",
              "source": {
                "line": 677,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L677"
              },
              "summary": "Floor of x as an int32t."
            },
            {
              "description": "Compute `2^n` as a float16 by building the exponent field directly.",
              "examples": [],
              "id": "arm_nn_softmax_exp2i_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_exp2i_f16",
              "params": [
                {
                  "description": "Integer exponent. Clamped to the normal float16 exponent range `[-14, 15]`.",
                  "direction": "in",
                  "name": "n",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`2^n` as a float16."
                }
              ],
              "signature": "static float16_t arm_nn_softmax_exp2i_f16(int32_t n)",
              "source": {
                "line": 690,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L690"
              },
              "summary": "Compute 2^n as a float16 by building the exponent field directly."
            },
            {
              "description": "Taylor/Estrin exp approximation for float16 softmax helpers.\n\nThe evaluation uses float32 intermediates to keep the approximation stable, but it is fully independent from the float32 softmax support tables.",
              "examples": [],
              "id": "arm_nn_softmax_exp_taylor_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_exp_taylor_f16",
              "params": [
                {
                  "description": "Exponent argument. Clamped to `[-80, 80]` before evaluation.",
                  "direction": "in",
                  "name": "x",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Approximation of `exp(x)`."
                }
              ],
              "signature": "static float16_t arm_nn_softmax_exp_taylor_f16(float16_t x)",
              "source": {
                "line": 710,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L710"
              },
              "summary": "Taylor/Estrin exp approximation for float16 softmax helpers."
            },
            {
              "description": "LUT-based exp approximation for float16 softmax helpers.\n\nSplits `x * log2(e)` into an integer part handled by arm_nn_softmax_exp2i_f16() and a fractional part interpolated linearly from `arm_nn_exp2_lut_f16`, using float32 intermediates.",
              "examples": [],
              "id": "arm_nn_softmax_exp_lut_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_exp_lut_f16",
              "params": [
                {
                  "description": "Exponent argument. Clamped to `[-80, 80]` before evaluation.",
                  "direction": "in",
                  "name": "x",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Approximation of `exp(x)`."
                }
              ],
              "signature": "static float16_t arm_nn_softmax_exp_lut_f16(float16_t x)",
              "source": {
                "line": 746,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L746"
              },
              "summary": "LUT-based exp approximation for float16 softmax helpers."
            },
            {
              "description": "Scalar exp approximation used by the float16 softmax paths.\n\nDispatches to arm_nn_softmax_exp_taylor_f16() when `ARM_NN_USE_EXP_TAYLOR` is defined and to arm_nn_softmax_exp_lut_f16() otherwise.",
              "examples": [],
              "id": "arm_nn_softmax_exp_scalar_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_exp_scalar_f16",
              "params": [
                {
                  "description": "Exponent argument.",
                  "direction": "in",
                  "name": "x",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Approximation of `exp(x)`."
                }
              ],
              "signature": "static float16_t arm_nn_softmax_exp_scalar_f16(float16_t x)",
              "source": {
                "line": 787,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L787"
              },
              "summary": "Scalar exp approximation used by the float16 softmax paths."
            },
            {
              "description": "Copy a float16 vector.",
              "examples": [],
              "id": "arm_memcpy_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memcpy_f16",
              "params": [
                {
                  "description": "Destination buffer.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float16_t *"
                },
                {
                  "description": "Source buffer.",
                  "direction": "in",
                  "name": "src",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of elements to copy.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memcpy_f16(float16_t *dst, const float16_t *src, uint32_t block_size)",
              "source": {
                "line": 979,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L979"
              },
              "summary": "Copy a float16 vector."
            },
            {
              "description": "Set a float16 vector to a constant value.",
              "examples": [],
              "id": "arm_memset_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memset_f16",
              "params": [
                {
                  "description": "Destination buffer.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float16_t *"
                },
                {
                  "description": "Fill value.",
                  "direction": "in",
                  "name": "val",
                  "type": "const float16_t"
                },
                {
                  "description": "Number of elements to write.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memset_f16(float16_t *dst, const float16_t val, uint32_t block_size)",
              "source": {
                "line": 1002,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1002"
              },
              "summary": "Set a float16 vector to a constant value."
            },
            {
              "description": "Specialized NHWC depthwise `2x5` kernel (float16).",
              "examples": [],
              "id": "arm_nn_depthwise_conv2x5_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv2x5_nhwc_f16",
              "params": [
                {
                  "description": "Input tensor in NHWC layout with shape `[batches][2][in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of batches.",
                  "direction": "in",
                  "name": "batches",
                  "type": "int32_t"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Channel multiplier; the output has `in_c * ch_mult` channels.",
                  "direction": "in",
                  "name": "ch_mult",
                  "type": "int32_t"
                },
                {
                  "description": "Depthwise weights with shape `[2][5][in_c * ch_mult]`.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `in_c * ch_mult` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor in NHWC layout with shape `[batches][1][out_w][in_c * ch_mult]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Output width. Output position `ow` reads input columns `ow..ow+4`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                },
                {
                  "description": "Lower clamp bound applied to `out`.",
                  "direction": "in",
                  "name": "act_min",
                  "type": "float16_t"
                },
                {
                  "description": "Upper clamp bound applied to `out`.",
                  "direction": "in",
                  "name": "act_max",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_depthwise_conv2x5_nhwc_f16(\n    const float16_t *x_nhwc,\n    int32_t batches,\n    int32_t in_c,\n    int32_t in_w,\n    int32_t ch_mult,\n    const float16_t *kernel,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_w,\n    float16_t act_min,\n    float16_t act_max\n)",
              "source": {
                "line": 1039,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1039"
              },
              "summary": "Specialized NHWC depthwise 2x5 kernel (float16)."
            },
            {
              "description": "Specialized NHWC depthwise 1D kernel for `k=3`, `ch_mult=1` (float32).",
              "examples": [],
              "id": "arm_nn_depthwise_conv1d_k3_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv1d_k3_nhwc_f16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of input (and output) channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Depthwise weights with shape `[3][in_c]`.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `in_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][in_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_depthwise_conv1d_k3_nhwc_f16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float16_t *kernel,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_w\n)",
              "source": {
                "line": 1054,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1054"
              },
              "summary": "Specialized NHWC depthwise 1D kernel for k=3, chmult=1 (float32)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=5` (float32).\n\n:::note\nMVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 5 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 6 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).\n\n:::",
              "examples": [],
              "id": "arm_nn_conv1d_k5_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k5_nhwc_f16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights with shape `[out_c][5][in_c]`.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+4`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k5_nhwc_f16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float16_t *kernel,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 1073,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1073"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=5 (float32)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=5` (float32).\n\n:::note\nMVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 5 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 6 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).\n\n:::\n\n:::note\nEvery MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k5_nhwc_f16.\n\n:::",
              "examples": [],
              "id": "arm_nn_conv1d_k5_nhwc_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k5_nhwc_f16_acc16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights with shape `[out_c][5][in_c]`.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+4`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k5_nhwc_f16_acc16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float16_t *kernel,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 1088,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1088"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=5 (float32)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=5` (float16, packed weights).\n\nThe packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(5 * in_c)][out_c_block_of_8]`.\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 6 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 6) this is the float16-lane result.\n\n:::",
              "examples": [],
              "id": "arm_nn_conv1d_k5_packed_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k5_packed_f16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights packed in output-channel blocks of 8 as described above.",
                  "direction": "in",
                  "name": "kernel_packed",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+4`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k5_packed_f16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float16_t *kernel_packed,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 1118,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1118"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=5 (float16, packed weights)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=5` (float16, packed weights).\n\nThe packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(5 * in_c)][out_c_block_of_8]`.\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 6 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 6) this is the float16-lane result.\n\n:::\n\n:::note\nEvery MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k5_packed_f16.\n\n:::",
              "examples": [],
              "id": "arm_nn_conv1d_k5_packed_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k5_packed_f16_acc16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights packed in output-channel blocks of 8 as described above.",
                  "direction": "in",
                  "name": "kernel_packed",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+4`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k5_packed_f16_acc16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float16_t *kernel_packed,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 1133,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1133"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=5 (float16, packed weights)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=3` (float32).\n\n:::note\nMVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 3 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 10 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).\n\n:::",
              "examples": [],
              "id": "arm_nn_conv1d_k3_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k3_nhwc_f16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights with shape `[out_c][3][in_c]`.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k3_nhwc_f16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float16_t *kernel,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 1153,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1153"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=3 (float32)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=3` (float32).\n\n:::note\nMVE leg: blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586). Input channel c feeds lane c % 8 with 3 taps per channel step; above 32 taps per output, a lane's float16 partial covers at most 10 channel steps; each block's lanes are folded into float32 pair accumulators (arm_nn_f16_fold_pairs_f32), which are summed once (arm_nn_f16_pairs_sum_f32), the bias is added in float32 and the total rounds to float16 once. Up to 32 taps: float16 lanes, a float16 reduction and the bias added in float16, as before, in the order the compiler gives them (it may reorder them under -ffast-math); only the fold's order is fixed. The scalar leg accumulates in float32 (#449, #465).\n\n:::\n\n:::note\nEvery MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k3_nhwc_f16.\n\n:::",
              "examples": [],
              "id": "arm_nn_conv1d_k3_nhwc_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k3_nhwc_f16_acc16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights with shape `[out_c][3][in_c]`.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k3_nhwc_f16_acc16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float16_t *kernel,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 1168,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1168"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=3 (float32)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=3` (float16, packed weights).\n\nThe packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(3 * in_c)][out_c_block_of_8]`.\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 10 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 10) this is the float16-lane result.\n\n:::",
              "examples": [],
              "id": "arm_nn_conv1d_k3_packed_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k3_packed_f16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights packed in output-channel blocks of 8 as described above.",
                  "direction": "in",
                  "name": "kernel_packed",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k3_packed_f16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float16_t *kernel_packed,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 1198,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1198"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=3 (float16, packed weights)."
            },
            {
              "description": "Specialized NHWC 1D convolution kernel for `k=3` (float16, packed weights).\n\nThe packed kernel uses the same `NTxN` RHS layout as `arm_nn_mat_mult_nt_n_packed_f16`, i.e. `[(3 * in_c)][out_c_block_of_8]`.\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and products accumulate in float32 and round to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: blockwise (#586): a lane's float16 partial covers at most 10 input channels (30 taps, bias first) before it is widened into a float32 accumulator; one rounding at the store. Up to 32 taps (in_c <= 10) this is the float16-lane result.\n\n:::\n\n:::note\nEvery MVE accumulator lane stays in float16 (no blockwise fold, AmbiqAI/ns-cmsis-nn#586); the scalar leg is the same as arm_nn_conv1d_k3_packed_f16.\n\n:::",
              "examples": [],
              "id": "arm_nn_conv1d_k3_packed_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_conv1d_k3_packed_f16_acc16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Weights packed in output-channel blocks of 8 as described above.",
                  "direction": "in",
                  "name": "kernel_packed",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `out_c` elements. May be NULL.",
                  "direction": "in",
                  "name": "b",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][out_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of output channels.",
                  "direction": "in",
                  "name": "out_c",
                  "type": "int32_t"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `ow..ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_conv1d_k3_packed_f16_acc16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    const float16_t *kernel_packed,\n    const float16_t *b,\n    float16_t *out,\n    int32_t out_c,\n    int32_t out_w\n)",
              "source": {
                "line": 1213,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1213"
              },
              "summary": "Specialized NHWC 1D convolution kernel for k=3 (float16, packed weights)."
            },
            {
              "description": "Specialized NHWC max-pool 1D kernel for `k=3`, `s=3` (float16).",
              "examples": [],
              "id": "arm_nn_maxpool1d_k3s3_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_maxpool1d_k3s3_nhwc_f16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][in_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `3*ow..3*ow+2`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_maxpool1d_k3s3_nhwc_f16(const float16_t *x_nhwc, int32_t in_c, int32_t in_w, float16_t *out, int32_t out_w)",
              "source": {
                "line": 1227,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1227"
              },
              "summary": "Specialized NHWC max-pool 1D kernel for k=3, s=3 (float16)."
            },
            {
              "description": "Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` without output clamp (float16).",
              "examples": [],
              "id": "arm_nn_maxpool1d_k2s2_nhwc_noclip_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_maxpool1d_k2s2_nhwc_noclip_f16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][in_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `2*ow..2*ow+1`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_maxpool1d_k2s2_nhwc_noclip_f16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    float16_t *out,\n    int32_t out_w\n)",
              "source": {
                "line": 1238,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1238"
              },
              "summary": "Specialized NHWC max-pool 1D kernel for k=2, s=2 without output clamp (float16)."
            },
            {
              "description": "Specialized NHWC max-pool 1D kernel for `k=2`, `s=2` with clamp (float16).",
              "examples": [],
              "id": "arm_nn_maxpool1d_k2s2_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_maxpool1d_k2s2_nhwc_f16",
              "params": [
                {
                  "description": "Input row in NHWC layout with shape `[in_w][in_c]`.",
                  "direction": "in",
                  "name": "x_nhwc",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Input width. Currently unused by the kernel.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Output row in NHWC layout with shape `[out_w][in_c]`.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                },
                {
                  "description": "Output width. Output position `ow` reads input positions `2*ow..2*ow+1`.",
                  "direction": "in",
                  "name": "out_w",
                  "type": "int32_t"
                },
                {
                  "description": "Lower clamp bound applied to `out`.",
                  "direction": "in",
                  "name": "act_min",
                  "type": "float16_t"
                },
                {
                  "description": "Upper clamp bound applied to `out`.",
                  "direction": "in",
                  "name": "act_max",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_maxpool1d_k2s2_nhwc_f16(\n    const float16_t *x_nhwc,\n    int32_t in_c,\n    int32_t in_w,\n    float16_t *out,\n    int32_t out_w,\n    float16_t act_min,\n    float16_t act_max\n)",
              "source": {
                "line": 1249,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1249"
              },
              "summary": "Specialized NHWC max-pool 1D kernel for k=2, s=2 with clamp (float16)."
            },
            {
              "description": "Matrix multiply with non-transposed lhs and transposed rhs rows (float32).\n\n:::note\nAccumulation width per leg. MVE legs accumulate blockwise (AmbiqAI/ns-cmsis-nn#586, superseding the float16-lane choice of #417 / #446 for the MVE legs). Up to rhs_cols 32 nothing changes: per-k float16 lanes on the gather path (rhs_cols below the contiguous-K threshold), lane-partial sums then one float16 reduction plus the bias in float16 at rhs_cols 32. Above 32, on the contiguous-K path and the remainder rows, each lane (element k goes to lane k % 8) sums at most 32 of its own taps (256 elements) in float16; each block's lanes are then widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (set by the first block, added to by later ones); the four pair accumulators are summed once as (0+1) + (2+3), so a single block sums ((0+1) + (2+3)) + ((4+5) + (6+7)), the bias is added in float32 and the total rounds to float16 once before the clamp. The float16 reduction up to rhs_cols 32 is ordered by the compiler, which may reorder it under -ffast-math; only the fold's order is fixed. arm_nn_mat_mult_nt_t_f16_acc16 keeps the float16 lanes and float16 reduction throughout. The scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates bias and every product in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457).\n\n:::",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_t_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_t_f16",
              "params": [
                {
                  "description": "Left-hand matrix stored row-major.",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const float16_t *"
                },
                {
                  "description": "Right-hand matrix stored row-major, one row per output channel.",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output matrix.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of rows in `lhs`.",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of rows in `rhs`.",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of columns in `rhs`.",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "int32_t"
                },
                {
                  "description": "Output row stride, expressed in elements.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "int32_t"
                },
                {
                  "description": "Lower clamp bound.",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Upper clamp bound.",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f16(\n    const float16_t *lhs,\n    const float16_t *rhs,\n    const float16_t *bias,\n    float16_t *dst,\n    int32_t lhs_rows,\n    int32_t rhs_rows,\n    int32_t rhs_cols,\n    int32_t row_address_offset,\n    float16_t activation_min,\n    float16_t activation_max\n)",
              "source": {
                "line": 1273,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1273"
              },
              "summary": "Matrix multiply with non-transposed lhs and transposed rhs rows (float32)."
            },
            {
              "description": "arm_nn_mat_mult_nt_t_f16 with every MVE accumulator lane in float16 (no blockwise fold).\n\nSame arguments, return codes and scalar leg as arm_nn_mat_mult_nt_t_f16; see its accumulation note.",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_t_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_t_f16_acc16",
              "params": [
                {
                  "description": "Left-hand matrix, row-major `[lhs_rows, rhs_cols]`.",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const float16_t *"
                },
                {
                  "description": "Right-hand matrix, row-major `[rhs_rows, rhs_cols]` (transposed operand).",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector of `rhs_rows` elements.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output matrix.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of rows in `lhs`.",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of rows in `rhs`.",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Shared reduction dimension `K`.",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "int32_t"
                },
                {
                  "description": "Output row stride, expressed in elements.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "int32_t"
                },
                {
                  "description": "Lower clamp bound.",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Upper clamp bound.",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_t_f16_acc16(\n    const float16_t *lhs,\n    const float16_t *rhs,\n    const float16_t *bias,\n    float16_t *dst,\n    int32_t lhs_rows,\n    int32_t rhs_rows,\n    int32_t rhs_cols,\n    int32_t row_address_offset,\n    float16_t activation_min,\n    float16_t activation_max\n)",
              "source": {
                "line": 1301,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1301"
              },
              "summary": "armnnmatmultnttf16 with every MVE accumulator lane in float16 (no blockwise fold)."
            },
            {
              "description": "Matrix multiply with non-transposed lhs and packed non-transposed rhs (float16).\n\n:::note\nOn non-MVE builds the output clamp is the bit-classified scalar clamp of #380, so a NaN accumulator (a NaN in `lhs`, `rhs_packed` or `bias`) propagates to `dst` at every optimization level on the gated toolchains, including the shipped -Ofast. On MVE builds the clamp is vmaxnmq/vminnmq with no NaN restore, so a NaN resolves to a clamp bound there instead.\n\n:::\n\n:::note\nAccumulation width per leg: the MVE leg accumulates blockwise (AmbiqAI/ns-cmsis-nn#586): one lane per output column, per-k, the bias opening the first block; every 32 k the float16 partial is widened exactly into per-lane float32 accumulators, which round to float16 once before the clamp (rhs_cols up to 32: exactly the float16-lane result; arm_nn_mat_mult_nt_n_packed_f16_acc16 keeps float16 lanes throughout). The scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE) accumulates bias and every product in float32 and rounds to float16 once before the clamp (AmbiqAI/ns-cmsis-nn#449, #457).\n\n:::",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_n_packed_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_n_packed_f16",
              "params": [
                {
                  "description": "Left-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`.",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const float16_t *"
                },
                {
                  "description": "Right-hand matrix with logical shape `[rhs_cols, rhs_rows]`, packed in column blocks of 8. The final block uses the same packed stride and inactive tail lanes are ignored.",
                  "direction": "in",
                  "name": "rhs_packed",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output matrix.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of rows in `lhs`.",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of logical output columns in the unpacked rhs matrix.",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Shared reduction dimension `K`.",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "int32_t"
                },
                {
                  "description": "Output row stride, expressed in elements.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "int32_t"
                },
                {
                  "description": "Lower clamp bound.",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Upper clamp bound.",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f16(\n    const float16_t *lhs,\n    const float16_t *rhs_packed,\n    const float16_t *bias,\n    float16_t *dst,\n    int32_t lhs_rows,\n    int32_t rhs_rows,\n    int32_t rhs_cols,\n    int32_t row_address_offset,\n    float16_t activation_min,\n    float16_t activation_max\n)",
              "source": {
                "line": 1342,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1342"
              },
              "summary": "Matrix multiply with non-transposed lhs and packed non-transposed rhs (float16)."
            },
            {
              "description": "arm_nn_mat_mult_nt_n_packed_f16 with every MVE accumulator lane in float16 (no blockwise fold).\n\nSame arguments, return codes and scalar leg as arm_nn_mat_mult_nt_n_packed_f16; see its accumulation note.",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_n_packed_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_n_packed_f16_acc16",
              "params": [
                {
                  "description": "Left-hand matrix stored row-major with logical shape `[lhs_rows, rhs_cols]`.",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const float16_t *"
                },
                {
                  "description": "Right-hand matrix packed in column blocks of 8.",
                  "direction": "in",
                  "name": "rhs_packed",
                  "type": "const float16_t *"
                },
                {
                  "description": "Optional bias vector.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output matrix.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of rows in `lhs`.",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of logical output columns in the unpacked rhs matrix.",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Shared reduction dimension `K`.",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "int32_t"
                },
                {
                  "description": "Output row stride, expressed in elements.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "int32_t"
                },
                {
                  "description": "Lower clamp bound.",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "float16_t"
                },
                {
                  "description": "Upper clamp bound.",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "float16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_n_packed_f16_acc16(\n    const float16_t *lhs,\n    const float16_t *rhs_packed,\n    const float16_t *bias,\n    float16_t *dst,\n    int32_t lhs_rows,\n    int32_t rhs_rows,\n    int32_t rhs_cols,\n    int32_t row_address_offset,\n    float16_t activation_min,\n    float16_t activation_max\n)",
              "source": {
                "line": 1370,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1370"
              },
              "summary": "armnnmatmultntnpackedf16 with every MVE accumulator lane in float16 (no blockwise fold)."
            },
            {
              "description": "Update LSTM function for an iteration step using float16 input, output and state.",
              "examples": [],
              "id": "arm_nn_lstm_step_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_lstm_step_f16",
              "params": [
                {
                  "description": "Data input pointer.",
                  "direction": "in",
                  "name": "data_in",
                  "type": "const float16_t *"
                },
                {
                  "description": "Hidden state / recurrent input pointer. May be NULL for the first step.",
                  "direction": "in",
                  "name": "hidden_in",
                  "type": "const float16_t *"
                },
                {
                  "description": "Hidden state / recurrent output pointer.",
                  "direction": "out",
                  "name": "hidden_out",
                  "type": "float16_t *"
                },
                {
                  "description": "Struct containing all information about the LSTM operator.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params_f16 *"
                },
                {
                  "description": "Struct containing pointers to mutable cell-state storage.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_lstm_context_f16 *"
                },
                {
                  "description": "Number of timesteps between consecutive batches.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params/buffers or buffers->cell_state, batch_offset <= 0)."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_lstm_step_f16(\n    const float16_t *data_in,\n    const float16_t *hidden_in,\n    float16_t *hidden_out,\n    const cmsis_nn_lstm_params_f16 *params,\n    cmsis_nn_lstm_context_f16 *buffers,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 1394,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1394"
              },
              "summary": "Update LSTM function for an iteration step using float16 input, output and state."
            },
            {
              "description": "Update GRU function for a single iteration step using float16 data.",
              "examples": [],
              "id": "arm_nn_gru_step_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_gru_step_f16",
              "params": [
                {
                  "description": "Data input pointer for this time step.",
                  "direction": "in",
                  "name": "data_in",
                  "type": "const float16_t *"
                },
                {
                  "description": "Recurrent input pointer. NULL for the first step (h_prev = 0).",
                  "direction": "in",
                  "name": "hidden_in",
                  "type": "const float16_t *"
                },
                {
                  "description": "Hidden-state output pointer for this time step.",
                  "direction": "out",
                  "name": "hidden_out",
                  "type": "float16_t *"
                },
                {
                  "description": "Struct describing the GRU operator.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gru_params_f16 *"
                },
                {
                  "description": "Scratch buffers. temp1 (>= hidden_size) is required when reset_after == 0.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_gru_context_f16 *"
                },
                {
                  "description": "Number of timesteps between consecutive batches.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params, batch_offset <= 0, or missing temp1 when reset_after == 0)."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_gru_step_f16(\n    const float16_t *data_in,\n    const float16_t *hidden_in,\n    float16_t *hidden_out,\n    const cmsis_nn_gru_params_f16 *params,\n    cmsis_nn_gru_context_f16 *buffers,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 1414,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1414"
              },
              "summary": "Update GRU function for a single iteration step using float16 data."
            },
            {
              "description": "Pack a single convolution patch into one row of a contiguous float32 patch matrix.\n\nDevelopers familiar with im2row/im2col terminology can think of this as packing one output patch into one row.",
              "examples": [],
              "id": "arm_nn_pack_conv_patch_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_pack_conv_patch_f16",
              "params": [
                {
                  "description": "Input tensor for one batch in NHWC layout with shape `[in_h][in_w][in_c]`.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Input height.",
                  "direction": "in",
                  "name": "in_h",
                  "type": "int32_t"
                },
                {
                  "description": "Input width.",
                  "direction": "in",
                  "name": "in_w",
                  "type": "int32_t"
                },
                {
                  "description": "Number of input channels.",
                  "direction": "in",
                  "name": "in_c",
                  "type": "int32_t"
                },
                {
                  "description": "Kernel height.",
                  "direction": "in",
                  "name": "kernel_h",
                  "type": "int32_t"
                },
                {
                  "description": "Kernel width.",
                  "direction": "in",
                  "name": "kernel_w",
                  "type": "int32_t"
                },
                {
                  "description": "Vertical stride.",
                  "direction": "in",
                  "name": "stride_h",
                  "type": "int32_t"
                },
                {
                  "description": "Horizontal stride.",
                  "direction": "in",
                  "name": "stride_w",
                  "type": "int32_t"
                },
                {
                  "description": "Top padding.",
                  "direction": "in",
                  "name": "pad_h",
                  "type": "int32_t"
                },
                {
                  "description": "Left padding.",
                  "direction": "in",
                  "name": "pad_w",
                  "type": "int32_t"
                },
                {
                  "description": "Vertical dilation.",
                  "direction": "in",
                  "name": "dilation_h",
                  "type": "int32_t"
                },
                {
                  "description": "Horizontal dilation.",
                  "direction": "in",
                  "name": "dilation_w",
                  "type": "int32_t"
                },
                {
                  "description": "Output row index of the patch to pack.",
                  "direction": "in",
                  "name": "out_y",
                  "type": "int32_t"
                },
                {
                  "description": "Output column index of the patch to pack.",
                  "direction": "in",
                  "name": "out_x",
                  "type": "int32_t"
                },
                {
                  "description": "Value written for taps that fall outside the input.",
                  "direction": "in",
                  "name": "pad_value",
                  "type": "float16_t"
                },
                {
                  "description": "Destination row of `kernel_h * kernel_w * in_c` elements, ordered `[kernel_h][kernel_w][in_c]`.",
                  "direction": "out",
                  "name": "patch_row",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_pack_conv_patch_f16(\n    const float16_t *input,\n    int32_t in_h,\n    int32_t in_w,\n    int32_t in_c,\n    int32_t kernel_h,\n    int32_t kernel_w,\n    int32_t stride_h,\n    int32_t stride_w,\n    int32_t pad_h,\n    int32_t pad_w,\n    int32_t dilation_h,\n    int32_t dilation_w,\n    int32_t out_y,\n    int32_t out_x,\n    float16_t pad_value,\n    float16_t *patch_row\n)",
              "source": {
                "line": 1424,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1424"
              },
              "summary": "Pack a single convolution patch into one row of a contiguous float32 patch matrix."
            },
            {
              "description": "Specialized softmax helper for a single float16 row of length 2.",
              "examples": [],
              "id": "arm_nn_softmax_1x2_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_1x2_f16",
              "params": [
                {
                  "description": "Pointer to two contiguous float16 input values.",
                  "direction": "in",
                  "name": "in",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to two contiguous float16 output values.",
                  "direction": "out",
                  "name": "out",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_softmax_1x2_f16(const float16_t *in, float16_t *out)",
              "source": {
                "line": 1447,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1447"
              },
              "summary": "Specialized softmax helper for a single float16 row of length 2."
            },
            {
              "description": "Update LSTM function for an iteration step using float32 input, output and state.",
              "examples": [],
              "id": "arm_nn_lstm_step_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_lstm_step_f32",
              "params": [
                {
                  "description": "Data input pointer.",
                  "direction": "in",
                  "name": "data_in",
                  "type": "const float32_t *"
                },
                {
                  "description": "Hidden state / recurrent input pointer. May be NULL for the first step.",
                  "direction": "in",
                  "name": "hidden_in",
                  "type": "const float32_t *"
                },
                {
                  "description": "Hidden state / recurrent output pointer.",
                  "direction": "out",
                  "name": "hidden_out",
                  "type": "float32_t *"
                },
                {
                  "description": "Struct containing all information about the LSTM operator.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params_f32 *"
                },
                {
                  "description": "Struct containing pointers to mutable cell-state storage.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_lstm_context_f32 *"
                },
                {
                  "description": "Number of timesteps between consecutive batches.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params/buffers or buffers->cell_state, batch_offset <= 0)."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_lstm_step_f32(\n    const float32_t *data_in,\n    const float32_t *hidden_in,\n    float32_t *hidden_out,\n    const cmsis_nn_lstm_params_f32 *params,\n    cmsis_nn_lstm_context_f32 *buffers,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 1466,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1466"
              },
              "summary": "Update LSTM function for an iteration step using float32 input, output and state."
            },
            {
              "description": "Update GRU function for a single iteration step using float32 data.",
              "examples": [],
              "id": "arm_nn_gru_step_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_gru_step_f32",
              "params": [
                {
                  "description": "Data input pointer for this time step.",
                  "direction": "in",
                  "name": "data_in",
                  "type": "const float32_t *"
                },
                {
                  "description": "Recurrent input pointer. NULL for the first step (h_prev = 0).",
                  "direction": "in",
                  "name": "hidden_in",
                  "type": "const float32_t *"
                },
                {
                  "description": "Hidden-state output pointer for this time step.",
                  "direction": "out",
                  "name": "hidden_out",
                  "type": "float32_t *"
                },
                {
                  "description": "Struct describing the GRU operator.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gru_params_f32 *"
                },
                {
                  "description": "Scratch buffers. temp1 (>= hidden_size) is required when reset_after == 0.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_gru_context_f32 *"
                },
                {
                  "description": "Number of timesteps between consecutive batches.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR on invalid arguments (NULL data_in/hidden_out/params, batch_offset <= 0, or missing temp1 when reset_after == 0)."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_gru_step_f32(\n    const float32_t *data_in,\n    const float32_t *hidden_in,\n    float32_t *hidden_out,\n    const cmsis_nn_gru_params_f32 *params,\n    cmsis_nn_gru_context_f32 *buffers,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 1486,
                "path": "Include/arm_nnsupportfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions_flt.h#L1486"
              },
              "summary": "Update GRU function for a single iteration step using float32 data."
            }
          ]
        },
        {
          "description": "",
          "name": "LSTM Layer Functions",
          "path": "heliaCORE.LSTM",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Unidirectional LSTM inference.\n\n:::note\nOn the MVE float path a NaN cell state yields a NaN hidden state, as tanh returns NaN unchanged there (#635). With cell clipping enabled the clip removes a NaN cell state first.\n\n:::",
              "examples": [],
              "id": "arm_lstm_unidirectional_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_f32",
              "params": [
                {
                  "description": "Pointer to the input sequence tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output sequence tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "LSTM parameters and weights.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params_f32 *"
                },
                {
                  "description": "Mutable LSTM scratch and state buffers. temp1 and temp2 are sized by `arm_lstm_unidirectional_f32_temp1_get_buffer_size()` / `arm_lstm_unidirectional_f32_temp2_get_buffer_size()`, which report 0: the float implementation never dereferences them and both may be NULL.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_lstm_context_f32 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_lstm_unidirectional_f32(\n    const float32_t *input,\n    float32_t *output,\n    const cmsis_nn_lstm_params_f32 *params,\n    cmsis_nn_lstm_context_f32 *buffers\n)",
              "source": {
                "line": 1781,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1781"
              },
              "summary": "Unidirectional LSTM inference."
            },
            {
              "description": "Unidirectional GRU layer for float32 input, output and state.\n\nImplements the reset-after GRU (Keras / TFLite default) when `params->reset_after` is non-zero, and the pre-reset variant otherwise. The hidden state is zero-initialised for the first time step, unless `buffers->hidden_state` is supplied for streaming state carry (`batch_size == 1`), in which case it seeds the initial state and receives the final hidden state on return.\n\n:::note\nNaN contract: a NaN in `input`, the previous hidden state, or the candidate gate's weight or bias reaches every output unit it feeds, on the scalar and MVE legs alike and at the shipped -Ofast: the MVE block re-establishes NaN after the table tanh with an integer-domain test that fast-math cannot elide (#251). A NaN confined to the update or reset gate's weight or bias does not reach the output: the scalar sigmoid maps NaN to 1.0 (see the note on arm_nn_sigmoid_scalar_f32 in `arm_nnsupportfunctions_flt.h`). NaN payloads and signs are not preserved on the MVE leg (default NaN, architectural). Inf follows the arithmetic.\n\n:::",
              "examples": [],
              "id": "arm_gru_unidirectional_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gru_unidirectional_f32",
              "params": [
                {
                  "description": "Input sequence tensor. Must not overlap `output`: earlier outputs are re-read as the recurrent state for later time steps, so aliasing corrupts silently.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output (hidden-state) sequence tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Struct describing the GRU operator.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gru_params_f32 *"
                },
                {
                  "description": "Scratch buffers. May be NULL when `reset_after` != 0. temp1 is sized by `arm_gru_unidirectional_f32_temp1_get_buffer_size()`.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_gru_context_f32 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, ARM_CMSIS_NN_ARG_ERROR otherwise."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gru_unidirectional_f32(\n    const float32_t *input,\n    float32_t *output,\n    const cmsis_nn_gru_params_f32 *params,\n    cmsis_nn_gru_context_f32 *buffers\n)",
              "source": {
                "line": 1812,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1812"
              },
              "summary": "Unidirectional GRU layer for float32 input, output and state."
            },
            {
              "description": "Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_f32()`.\n\n:::note\nThis query reports its invalid input as -1, following the integer LSTM temp sizers (`arm_lstm_unidirectional_s8_temp1_get_buffer_size()`), not the 0 used by the float convolution and fully-connected queries in this header.\n\n:::",
              "examples": [],
              "id": "arm_lstm_unidirectional_f32_temp1_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_f32_temp1_get_buffer_size",
              "params": [
                {
                  "description": "LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f32` passed to `arm_lstm_unidirectional_f32()`. No field is read.",
                  "direction": "in",
                  "name": "lstm_params",
                  "type": "const cmsis_nn_lstm_params_f32 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "0 for any non-NULL lstm_params, on every build target: the float32 implementation computes its gate values per hidden unit in automatics and never dereferences temp1 or temp2, so both context pointers may be NULL. Returns -1 only for a NULL lstm_params. The query exists so arena-sizing code can treat every LSTM variant alike; a future implementation that starts staging gate vectors would change this figure, so size from the query rather than hard-coding 0."
                }
              ],
              "signature": "int32_t arm_lstm_unidirectional_f32_temp1_get_buffer_size(const cmsis_nn_lstm_params_f32 *lstm_params)",
              "source": {
                "line": 1833,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1833"
              },
              "summary": "Get size of the temp1 scratch buffer required by armlstmunidirectionalf32()."
            },
            {
              "description": "Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_f32()`. The contract is identical to `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, and the answer is the same 0 (temp2 is likewise never dereferenced).\n\n:::note\nThis query reports its invalid input as -1, following the integer LSTM temp sizers (`arm_lstm_unidirectional_s8_temp1_get_buffer_size()`), not the 0 used by the float convolution and fully-connected queries in this header.\n\n:::",
              "examples": [],
              "id": "arm_lstm_unidirectional_f32_temp2_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_f32_temp2_get_buffer_size",
              "params": [
                {
                  "description": "LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f32` passed to `arm_lstm_unidirectional_f32()`. No field is read.",
                  "direction": "in",
                  "name": "lstm_params",
                  "type": "const cmsis_nn_lstm_params_f32 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "0 for any non-NULL lstm_params, on every build target: the float32 implementation computes its gate values per hidden unit in automatics and never dereferences temp1 or temp2, so both context pointers may be NULL. Returns -1 only for a NULL lstm_params. The query exists so arena-sizing code can treat every LSTM variant alike; a future implementation that starts staging gate vectors would change this figure, so size from the query rather than hard-coding 0."
                }
              ],
              "signature": "int32_t arm_lstm_unidirectional_f32_temp2_get_buffer_size(const cmsis_nn_lstm_params_f32 *lstm_params)",
              "source": {
                "line": 1842,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1842"
              },
              "summary": "Get size of the temp2 scratch buffer required by armlstmunidirectionalf32()."
            },
            {
              "description": "Get size of the temp1 scratch buffer required by `arm_gru_unidirectional_f32()`.\n\n:::note\nOn the pre-reset path a 0 is only returned for the degenerate hidden_size == 0, which `arm_gru_unidirectional_f32()` rejects with ARM_CMSIS_NN_ARG_ERROR before any buffer access - so a 0 there never corresponds to a runnable call.\n\n:::\n\n:::note\nThis query reports an out-of-range shape as -1, following the integer LSTM temp sizers, not the 0 used by the float convolution and fully-connected queries in this header.\n\n:::",
              "examples": [],
              "id": "arm_gru_unidirectional_f32_temp1_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gru_unidirectional_f32_temp1_get_buffer_size",
              "params": [
                {
                  "description": "GRU operator parameters, i.e. the same `cmsis_nn_gru_params_f32` passed to `arm_gru_unidirectional_f32()`. Only reset_after and hidden_size are read.",
                  "direction": "in",
                  "name": "gru_params",
                  "type": "const cmsis_nn_gru_params_f32 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: hidden_size * sizeof(float32_t) when reset_after == 0 (the pre-reset formulation stages the r . h_prev vector in temp1; the vector is reused across batches and time steps, so neither batch_size nor time_steps enters), and 0 when reset_after != 0 (temp1 is never dereferenced and may be NULL). Returns -1 if gru_params is NULL, if hidden_size is negative, or if the byte count would not fit in an int32_t. The figure and the range checks are the same on every build target."
                }
              ],
              "signature": "int32_t arm_gru_unidirectional_f32_temp1_get_buffer_size(const cmsis_nn_gru_params_f32 *gru_params)",
              "source": {
                "line": 1863,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1863"
              },
              "summary": "Get size of the temp1 scratch buffer required by armgruunidirectionalf32()."
            },
            {
              "description": "Unidirectional LSTM inference.\n\n:::note\nOn the MVE float path a NaN cell state yields a NaN hidden state, as tanh returns NaN unchanged there (#635). With cell clipping enabled the clip removes a NaN cell state first.\n\n:::",
              "examples": [],
              "id": "arm_lstm_unidirectional_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_f16",
              "params": [
                {
                  "description": "Pointer to the input sequence tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output sequence tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "LSTM parameters and weights.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params_f16 *"
                },
                {
                  "description": "Mutable LSTM scratch and state buffers. temp1 and temp2 are sized by `arm_lstm_unidirectional_f32_temp1_get_buffer_size()` / `arm_lstm_unidirectional_f32_temp2_get_buffer_size()`, which report 0: the float implementation never dereferences them and both may be NULL.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_lstm_context_f16 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_lstm_unidirectional_f16(\n    const float16_t *input,\n    float16_t *output,\n    const cmsis_nn_lstm_params_f16 *params,\n    cmsis_nn_lstm_context_f16 *buffers\n)",
              "source": {
                "line": 3558,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3558"
              },
              "summary": "Unidirectional LSTM inference."
            },
            {
              "description": "Unidirectional GRU layer for float16 input, output and state.\n\nImplements the reset-after GRU (Keras / TFLite default) when `params->reset_after` is non-zero, and the pre-reset variant otherwise. The hidden state is zero-initialised for the first time step, unless `buffers->hidden_state` is supplied for streaming state carry (`batch_size == 1`), in which case it seeds the initial state and receives the final hidden state on return.\n\n:::note\nNaN contract: a NaN in `input`, the previous hidden state, or the candidate gate's weight or bias reaches every output unit it feeds, on the scalar and MVE legs alike and at the shipped -Ofast: the MVE block re-establishes NaN after the table tanh with an integer-domain test that fast-math cannot elide (#251). A NaN confined to the update or reset gate's weight or bias does not reach the output: the scalar sigmoid maps NaN to 1.0 (see the note on arm_nn_sigmoid_scalar_f32 in `arm_nnsupportfunctions_flt.h`). NaN payloads and signs are not preserved on the MVE leg (default NaN, architectural). Inf follows the arithmetic.\n\n:::",
              "examples": [],
              "id": "arm_gru_unidirectional_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gru_unidirectional_f16",
              "params": [
                {
                  "description": "Input sequence tensor. Must not overlap `output`: earlier outputs are re-read as the recurrent state for later time steps, so aliasing corrupts silently.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output (hidden-state) sequence tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Struct describing the GRU operator.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gru_params_f16 *"
                },
                {
                  "description": "Scratch buffers. May be NULL when `reset_after` != 0. temp1 is sized by `arm_gru_unidirectional_f16_temp1_get_buffer_size()`.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_gru_context_f16 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, ARM_CMSIS_NN_ARG_ERROR otherwise."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gru_unidirectional_f16(\n    const float16_t *input,\n    float16_t *output,\n    const cmsis_nn_gru_params_f16 *params,\n    cmsis_nn_gru_context_f16 *buffers\n)",
              "source": {
                "line": 3589,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3589"
              },
              "summary": "Unidirectional GRU layer for float16 input, output and state."
            },
            {
              "description": "Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_f16()`. The contract is identical to `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, and the answer is the same 0 on every build target (the float16 implementation likewise never dereferences temp1 or temp2, which may both be NULL).",
              "examples": [],
              "id": "arm_lstm_unidirectional_f16_temp1_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_f16_temp1_get_buffer_size",
              "params": [
                {
                  "description": "LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f16` passed to `arm_lstm_unidirectional_f16()`. No field is read.",
                  "direction": "in",
                  "name": "lstm_params",
                  "type": "const cmsis_nn_lstm_params_f16 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "0 for any non-NULL lstm_params, -1 for a NULL lstm_params."
                }
              ],
              "signature": "int32_t arm_lstm_unidirectional_f16_temp1_get_buffer_size(const cmsis_nn_lstm_params_f16 *lstm_params)",
              "source": {
                "line": 3605,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3605"
              },
              "summary": "Get size of the temp1 scratch buffer required by armlstmunidirectionalf16()."
            },
            {
              "description": "Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_f16()`. The contract is identical to `arm_lstm_unidirectional_f32_temp1_get_buffer_size()`, and the answer is the same 0.",
              "examples": [],
              "id": "arm_lstm_unidirectional_f16_temp2_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_f16_temp2_get_buffer_size",
              "params": [
                {
                  "description": "LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params_f16` passed to `arm_lstm_unidirectional_f16()`. No field is read.",
                  "direction": "in",
                  "name": "lstm_params",
                  "type": "const cmsis_nn_lstm_params_f16 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "0 for any non-NULL lstm_params, -1 for a NULL lstm_params."
                }
              ],
              "signature": "int32_t arm_lstm_unidirectional_f16_temp2_get_buffer_size(const cmsis_nn_lstm_params_f16 *lstm_params)",
              "source": {
                "line": 3613,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3613"
              },
              "summary": "Get size of the temp2 scratch buffer required by armlstmunidirectionalf16()."
            },
            {
              "description": "Get size of the temp1 scratch buffer required by `arm_gru_unidirectional_f16()`. See `arm_gru_unidirectional_f32_temp1_get_buffer_size()` for the -1-on-invalid contract and the pre-reset degenerate-0 note; both apply here unchanged.",
              "examples": [],
              "id": "arm_gru_unidirectional_f16_temp1_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gru_unidirectional_f16_temp1_get_buffer_size",
              "params": [
                {
                  "description": "GRU operator parameters, i.e. the same `cmsis_nn_gru_params_f16` passed to `arm_gru_unidirectional_f16()`. Only reset_after and hidden_size are read.",
                  "direction": "in",
                  "name": "gru_params",
                  "type": "const cmsis_nn_gru_params_f16 *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: hidden_size * sizeof(float16_t) when reset_after == 0, 0 when reset_after != 0 (temp1 is never dereferenced and may be NULL). Half the figure `arm_gru_unidirectional_f32_temp1_get_buffer_size()` returns for the same shape - sizing an f16 layer with the f32 query over-allocates, and the reverse under-allocates."
                }
              ],
              "signature": "int32_t arm_gru_unidirectional_f16_temp1_get_buffer_size(const cmsis_nn_gru_params_f16 *gru_params)",
              "source": {
                "line": 3628,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3628"
              },
              "summary": "Get size of the temp1 scratch buffer required by armgruunidirectionalf16()."
            }
          ]
        },
        {
          "description": "Collection of convolution, depthwise convolution functions and their variants.\n\nThe convolution is implemented in 2 steps: im2col and General Matrix Multiplication(GEMM)\n\nim2col is a process of converting each patch of image data into a column. After im2col, the convolution is computed as matrix-matrix multiplication.\n\nTo reduce the memory footprint, the im2col is performed partially. Each iteration, only a few column (i.e., patches) are generated followed by GEMM.",
          "name": "Convolution Functions",
          "path": "heliaCORE.NNConv",
          "submodules": [],
          "summary": "Collection of convolution, depthwise convolution functions and their variants.",
          "symbols": [
            {
              "description": "Depthwise convolution, NHWC layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_nhwc_conv_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_nhwc_conv_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions in NHWC format.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions in NHWC-compatible depthwise format.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions in NHWC format.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_nhwc_conv_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output\n)",
              "source": {
                "line": 79,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L79"
              },
              "summary": "Depthwise convolution, NHWC layout."
            },
            {
              "description": "Depthwise convolution, dispatch by layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 119,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L119"
              },
              "summary": "Depthwise convolution, dispatch by layout."
            },
            {
              "description": "Depthwise convolution wrapper using the CMSIS-NN baseline path.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters.",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_wrapper_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output\n)",
              "source": {
                "line": 158,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L158"
              },
              "summary": "Depthwise convolution wrapper using the CMSIS-NN baseline path."
            },
            {
              "description": "Get the temporary buffer size required by depthwise convolution.\n\n:::note\nOnly one route reads scratch: on MVE builds, an NHWC depthwise with ch_mult != 1, a single input channel and at least CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD output channels can run as a regular convolution, which needs the repacked filter, `ROUND_UP(output_dims->c, 4) * filter_dims->h * filter_dims->w * sizeof(float32_t)` bytes (`ROUND_UP(output_dims->c, 8)` and `sizeof(float16_t)` for `_f16`), plus `arm_convolve_wrapper_f32_get_buffer_size` (`_f16`) for that convolution. The query reserves that for every such layer; one an exact-shape specialization takes first runs without it, so the size is an upper bound there. Every other layer  the `ch_mult == 1` direct kernel and the generic kernel  runs without scratch and the query returns 0 (AmbiqAI/ns-cmsis-nn#448, #625).\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_f32_get_buffer_size",
              "params": [
                {
                  "description": "Depthwise convolution parameters.",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_f32_get_buffer_size(\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 189,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L189"
              },
              "summary": "Get the temporary buffer size required by depthwise convolution."
            },
            {
              "description": "Get the buffer size required by the depthwise convolution wrapper.",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_f32_get_buffer_size",
              "params": [
                {
                  "description": "Depthwise convolution parameters.",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_f32_get_buffer_size(\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 205,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L205"
              },
              "summary": "Get the buffer size required by the depthwise convolution wrapper."
            },
            {
              "description": "Convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
              "examples": [],
              "id": "arm_convolve_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_nhwc_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_nhwc_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
              "source": {
                "line": 232,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L232"
              },
              "summary": "Convolution, NHWC layout."
            },
            {
              "description": "Convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
              "examples": [],
              "id": "arm_convolve_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 266,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L266"
              },
              "summary": "Convolution, dispatch by layout."
            },
            {
              "description": "Convolution wrapper using the CMSIS-NN baseline path.",
              "examples": [],
              "id": "arm_convolve_wrapper_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_wrapper_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
              "source": {
                "line": 294,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L294"
              },
              "summary": "Convolution wrapper using the CMSIS-NN baseline path."
            },
            {
              "description": "1x1 convolution, NHWC layout.",
              "examples": [],
              "id": "arm_convolve_1x1_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_nhwc_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_nhwc_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
              "source": {
                "line": 321,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L321"
              },
              "summary": "1x1 convolution, NHWC layout."
            },
            {
              "description": "1x1 convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1x1_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 354,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L354"
              },
              "summary": "1x1 convolution, dispatch by layout."
            },
            {
              "description": "1xN convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1_x_n_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_nhwc_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
              "source": {
                "line": 391,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L391"
              },
              "summary": "1xN convolution, NHWC layout."
            },
            {
              "description": "1xN convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1_x_n_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 428,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L428"
              },
              "summary": "1xN convolution, dispatch by layout."
            },
            {
              "description": "Get the temporary buffer size required by convolution.\n\n:::note\nWhen `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still reports only the temporary input/im2col scratch requirement. Any offline-packed filter storage is expected to be provided by the caller.\n\n:::",
              "examples": [],
              "id": "arm_convolve_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_f32_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters.",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_convolve_f32_get_buffer_size(\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 456,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L456"
              },
              "summary": "Get the temporary buffer size required by convolution."
            },
            {
              "description": "Get the buffer size required by the convolution wrapper.",
              "examples": [],
              "id": "arm_convolve_wrapper_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_f32_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters.",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_f32_get_buffer_size(\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 472,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L472"
              },
              "summary": "Get the buffer size required by the convolution wrapper."
            },
            {
              "description": "Get the buffer size required by 1x1 convolution.\n\n:::note\nReturns `0` for the unity-stride no-pack path. For non-unity-stride NHWC 1x1 convolution, the returned scratch size enables the packed-tile + GEMM path. When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still excludes the offline-packed filter storage itself.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1x1_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_f32_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters.",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_convolve_1x1_f32_get_buffer_size(\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 493,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L493"
              },
              "summary": "Get the buffer size required by 1x1 convolution."
            },
            {
              "description": "Get the buffer size required by 1xN convolution.",
              "examples": [],
              "id": "arm_convolve_1_x_n_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_f32_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters.",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_convolve_1_x_n_f32_get_buffer_size(\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 510,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L510"
              },
              "summary": "Get the buffer size required by 1xN convolution."
            },
            {
              "description": "Transpose convolution wrapper using the CMSIS-NN baseline path.",
              "examples": [],
              "id": "arm_transpose_conv_wrapper_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_wrapper_f32",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Output context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_conv_wrapper_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 1540,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1540"
              },
              "summary": "Transpose convolution wrapper using the CMSIS-NN baseline path."
            },
            {
              "description": "Transpose convolution, NHWC layout.",
              "examples": [],
              "id": "arm_transpose_conv_nhwc_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_nhwc_f32",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Output context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_conv_nhwc_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
              "source": {
                "line": 1570,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1570"
              },
              "summary": "Transpose convolution, NHWC layout."
            },
            {
              "description": "Transpose convolution, dispatch by layout.",
              "examples": [],
              "id": "arm_transpose_conv_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_f32",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Output context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_conv_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 1600,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1600"
              },
              "summary": "Transpose convolution, dispatch by layout."
            },
            {
              "description": "Get the temporary buffer size required by transpose convolution.",
              "examples": [],
              "id": "arm_transpose_conv_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_f32_get_buffer_size",
              "params": [
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "out_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_transpose_conv_f32_get_buffer_size(\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *out_dims\n)",
              "source": {
                "line": 1623,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1623"
              },
              "summary": "Get the temporary buffer size required by transpose convolution."
            },
            {
              "description": "Get the reverse-convolution workspace size used by transpose convolution helpers.",
              "examples": [],
              "id": "arm_transpose_conv_f32_get_reverse_conv_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_f32_get_reverse_conv_buffer_size",
              "params": [
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no reverse-convolution buffer is needed."
                }
              ],
              "signature": "int32_t arm_transpose_conv_f32_get_reverse_conv_buffer_size(\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims\n)",
              "source": {
                "line": 1638,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1638"
              },
              "summary": "Get the reverse-convolution workspace size used by transpose convolution helpers."
            },
            {
              "description": "Depthwise convolution, NHWC layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_nhwc_conv_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_nhwc_conv_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions in NHWC format.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions in NHWC-compatible depthwise format.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions in NHWC format.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_nhwc_conv_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
              "source": {
                "line": 2224,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2224"
              },
              "summary": "Depthwise convolution, NHWC layout."
            },
            {
              "description": "Depthwise convolution, NHWC layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_nhwc_conv_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_nhwc_conv_f16.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_nhwc_conv_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_nhwc_conv_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions in NHWC format.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions in NHWC-compatible depthwise format.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions in NHWC format.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_nhwc_conv_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
              "source": {
                "line": 2243,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2243"
              },
              "summary": "Depthwise convolution, NHWC layout."
            },
            {
              "description": "Depthwise convolution, dispatch by layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2273,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2273"
              },
              "summary": "Depthwise convolution, dispatch by layout."
            },
            {
              "description": "Depthwise convolution, dispatch by layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_conv_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_conv_f16.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2293,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2293"
              },
              "summary": "Depthwise convolution, dispatch by layout."
            },
            {
              "description": "Depthwise convolution wrapper using the CMSIS-NN baseline path.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters.",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_wrapper_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
              "source": {
                "line": 2324,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2324"
              },
              "summary": "Depthwise convolution wrapper using the CMSIS-NN baseline path."
            },
            {
              "description": "Depthwise convolution wrapper using the CMSIS-NN baseline path.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_conv_wrapper_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_conv_wrapper_f16.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters.",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_wrapper_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
              "source": {
                "line": 2343,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2343"
              },
              "summary": "Depthwise convolution wrapper using the CMSIS-NN baseline path."
            },
            {
              "description": "Get the temporary buffer size required by depthwise convolution.\n\n:::note\nOnly one route reads scratch: on MVE builds, an NHWC depthwise with ch_mult != 1, a single input channel and at least CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD output channels can run as a regular convolution, which needs the repacked filter, `ROUND_UP(output_dims->c, 4) * filter_dims->h * filter_dims->w * sizeof(float32_t)` bytes (`ROUND_UP(output_dims->c, 8)` and `sizeof(float16_t)` for `_f16`), plus `arm_convolve_wrapper_f32_get_buffer_size` (`_f16`) for that convolution. The query reserves that for every such layer; one an exact-shape specialization takes first runs without it, so the size is an upper bound there. Every other layer  the `ch_mult == 1` direct kernel and the generic kernel  runs without scratch and the query returns 0 (AmbiqAI/ns-cmsis-nn#448, #625).\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_f16_get_buffer_size",
              "params": [
                {
                  "description": "Depthwise convolution parameters.",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_f16_get_buffer_size(\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2357,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2357"
              },
              "summary": "Get the temporary buffer size required by depthwise convolution."
            },
            {
              "description": "Get the buffer size required by the depthwise convolution wrapper.",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_f16_get_buffer_size",
              "params": [
                {
                  "description": "Depthwise convolution parameters.",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_f16_get_buffer_size(\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2366,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2366"
              },
              "summary": "Get the buffer size required by the depthwise convolution wrapper."
            },
            {
              "description": "Convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
              "examples": [],
              "id": "arm_convolve_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_nhwc_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_nhwc_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 2374,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2374"
              },
              "summary": "Convolution, NHWC layout."
            },
            {
              "description": "Convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_nhwc_f16.\n\n:::",
              "examples": [],
              "id": "arm_convolve_nhwc_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_nhwc_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_nhwc_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 2393,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2393"
              },
              "summary": "Convolution, NHWC layout."
            },
            {
              "description": "Convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct OHWI / NT_N_PACKED fallback accumulates bias and every tap in float32 and rounds to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #457); the 1x1, 1xN and patch-GEMM paths go through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16, whose scalar legs do the same, as do the 1xN no-padding OHWI region and the k=3 / k=5 conv1d specializations (#465). MVE leg: the direct small-C kernel accumulates in float32 (widened lanes); the direct OHWI / NT_N_PACKED fallback, every matmul-backed path (1x1, 1xN, patch-GEMM), the 1xN no-padding region and the conv1d specializations use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. Where a dot product spreads its taps over the lanes of one vector (OHWI rows, the contiguous-K matmul, the conv1d k=3 / k=5 OHWI kernels), each lane's own taps form its blocks and, once the reduction exceeds 32 taps, each block's lanes are widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (the first block sets it); the four pair accumulators are summed once as (0+1) + (2+3), the bias is added in float32 and the total rounds once. The k=3 / k=5 kernels close a block on a whole input-channel step (30 taps per lane). An output's taps are the ones its kernel multiplies: the direct fallback skips padded taps, so an edge output counts only its in-range taps, while patch-GEMM and the 1xN padded regions multiply a zero-padded patch and count its padded taps too. Patch-GEMM runs only when ctx provides its scratch, so an edge output's value can depend on whether ctx->buf is given. The fold's order is fixed; the float16 reduction of a dot of at most 32 taps is left to the compiler, which may reorder it under -ffast-math, as before #586.\n\n:::",
              "examples": [],
              "id": "arm_convolve_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2431,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2431"
              },
              "summary": "Convolution, dispatch by layout."
            },
            {
              "description": "Convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct OHWI / NT_N_PACKED fallback accumulates bias and every tap in float32 and rounds to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #457); the 1x1, 1xN and patch-GEMM paths go through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16, whose scalar legs do the same, as do the 1xN no-padding OHWI region and the k=3 / k=5 conv1d specializations (#465). MVE leg: the direct small-C kernel accumulates in float32 (widened lanes); the direct OHWI / NT_N_PACKED fallback, every matmul-backed path (1x1, 1xN, patch-GEMM), the 1xN no-padding region and the conv1d specializations use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. Where a dot product spreads its taps over the lanes of one vector (OHWI rows, the contiguous-K matmul, the conv1d k=3 / k=5 OHWI kernels), each lane's own taps form its blocks and, once the reduction exceeds 32 taps, each block's lanes are widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (the first block sets it); the four pair accumulators are summed once as (0+1) + (2+3), the bias is added in float32 and the total rounds once. The k=3 / k=5 kernels close a block on a whole input-channel step (30 taps per lane). An output's taps are the ones its kernel multiplies: the direct fallback skips padded taps, so an edge output counts only its in-range taps, while patch-GEMM and the 1xN padded regions multiply a zero-padded patch and count its padded taps too. Patch-GEMM runs only when ctx provides its scratch, so an edge output's value can depend on whether ctx->buf is given. The fold's order is fixed; the float16 reduction of a dot of at most 32 taps is left to the compiler, which may reorder it under -ffast-math, as before #586.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_f16.\n\n:::",
              "examples": [],
              "id": "arm_convolve_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2451,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2451"
              },
              "summary": "Convolution, dispatch by layout."
            },
            {
              "description": "Convolution wrapper using the CMSIS-NN baseline path.",
              "examples": [],
              "id": "arm_convolve_wrapper_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_wrapper_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 2466,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2466"
              },
              "summary": "Convolution wrapper using the CMSIS-NN baseline path."
            },
            {
              "description": "Convolution wrapper using the CMSIS-NN baseline path.\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_wrapper_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_wrapper_f16.\n\n:::",
              "examples": [],
              "id": "arm_convolve_wrapper_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_wrapper_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 2485,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2485"
              },
              "summary": "Convolution wrapper using the CMSIS-NN baseline path."
            },
            {
              "description": "1x1 convolution, NHWC layout.",
              "examples": [],
              "id": "arm_convolve_1x1_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_nhwc_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_nhwc_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 2499,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2499"
              },
              "summary": "1x1 convolution, NHWC layout."
            },
            {
              "description": "1x1 convolution, NHWC layout.\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1x1_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1x1_nhwc_f16.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1x1_nhwc_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_nhwc_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_nhwc_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 2518,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2518"
              },
              "summary": "1x1 convolution, NHWC layout."
            },
            {
              "description": "1x1 convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1x1_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2532,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2532"
              },
              "summary": "1x1 convolution, dispatch by layout."
            },
            {
              "description": "1x1 convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1x1_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1x1_f16.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1x1_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2552,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2552"
              },
              "summary": "1x1 convolution, dispatch by layout."
            },
            {
              "description": "1xN convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1_x_n_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_nhwc_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 2567,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2567"
              },
              "summary": "1xN convolution, NHWC layout."
            },
            {
              "description": "1xN convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1_x_n_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1_x_n_nhwc_f16.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1_x_n_nhwc_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_nhwc_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 2586,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2586"
              },
              "summary": "1xN convolution, NHWC layout."
            },
            {
              "description": "1xN convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and every product accumulate in float32 and round to float16 once (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: the padded regions go through the matmul helpers and the no-padding region through a strided kernel, all with blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. The padded regions multiply a zero-padded patch row, so their outputs count the padded taps as well.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1_x_n_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2610,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2610"
              },
              "summary": "1xN convolution, dispatch by layout."
            },
            {
              "description": "1xN convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and every product accumulate in float32 and round to float16 once (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: the padded regions go through the matmul helpers and the no-padding region through a strided kernel, all with blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. The padded regions multiply a zero-padded patch row, so their outputs count the padded taps as well.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1_x_n_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1_x_n_f16.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1_x_n_f16_acc16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_f16_acc16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT].",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format depends on `layout`.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2630,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2630"
              },
              "summary": "1xN convolution, dispatch by layout."
            },
            {
              "description": "Get the temporary buffer size required by convolution.\n\n:::note\nWhen `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still reports only the temporary input/im2col scratch requirement. Any offline-packed filter storage is expected to be provided by the caller.\n\n:::",
              "examples": [],
              "id": "arm_convolve_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_f16_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters.",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_convolve_f16_get_buffer_size(\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2645,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2645"
              },
              "summary": "Get the temporary buffer size required by convolution."
            },
            {
              "description": "Get the buffer size required by the convolution wrapper.",
              "examples": [],
              "id": "arm_convolve_wrapper_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_f16_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters.",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_f16_get_buffer_size(\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2654,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2654"
              },
              "summary": "Get the buffer size required by the convolution wrapper."
            },
            {
              "description": "Get the buffer size required by 1x1 convolution.\n\n:::note\nReturns `0` for the unity-stride no-pack path. For non-unity-stride NHWC 1x1 convolution, the returned scratch size enables the packed-tile + GEMM path. When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still excludes the offline-packed filter storage itself.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1x1_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_f16_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters.",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_convolve_1x1_f16_get_buffer_size(\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2662,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2662"
              },
              "summary": "Get the buffer size required by 1x1 convolution."
            },
            {
              "description": "Get the buffer size required by 1xN convolution.",
              "examples": [],
              "id": "arm_convolve_1_x_n_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_f16_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters.",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_convolve_1_x_n_f16_get_buffer_size(\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 2671,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2671"
              },
              "summary": "Get the buffer size required by 1xN convolution."
            },
            {
              "description": "Transpose convolution wrapper using the CMSIS-NN baseline path.",
              "examples": [],
              "id": "arm_transpose_conv_wrapper_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_wrapper_f16",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Output context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_conv_wrapper_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 3354,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3354"
              },
              "summary": "Transpose convolution wrapper using the CMSIS-NN baseline path."
            },
            {
              "description": "Transpose convolution, NHWC layout.",
              "examples": [],
              "id": "arm_transpose_conv_nhwc_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_nhwc_f16",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Output context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_conv_nhwc_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 3370,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3370"
              },
              "summary": "Transpose convolution, NHWC layout."
            },
            {
              "description": "Transpose convolution, dispatch by layout.",
              "examples": [],
              "id": "arm_transpose_conv_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_f16",
              "params": [
                {
                  "description": "Function context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Output context. Unused; may be NULL.",
                  "direction": "in",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the filter tensor data.",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_conv_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 3385,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3385"
              },
              "summary": "Transpose convolution, dispatch by layout."
            },
            {
              "description": "Get the temporary buffer size required by transpose convolution.",
              "examples": [],
              "id": "arm_transpose_conv_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_f16_get_buffer_size",
              "params": [
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "out_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
                }
              ],
              "signature": "int32_t arm_transpose_conv_f16_get_buffer_size(\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *out_dims\n)",
              "source": {
                "line": 3401,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3401"
              },
              "summary": "Get the temporary buffer size required by transpose convolution."
            },
            {
              "description": "Get the reverse-convolution workspace size used by transpose convolution helpers.",
              "examples": [],
              "id": "arm_transpose_conv_f16_get_reverse_conv_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_f16_get_reverse_conv_buffer_size",
              "params": [
                {
                  "description": "Transpose convolution parameters.",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes, or 0 when no reverse-convolution buffer is needed."
                }
              ],
              "signature": "int32_t arm_transpose_conv_f16_get_reverse_conv_buffer_size(\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims\n)",
              "source": {
                "line": 3410,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3410"
              },
              "summary": "Get the reverse-convolution workspace size used by transpose convolution helpers."
            }
          ]
        },
        {
          "description": "",
          "name": "NNSupport",
          "path": "heliaCORE.NNSupport",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Transpose a floating-point tensor.",
              "examples": [],
              "id": "arm_transpose_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Transpose parameters, including permutation and layout information. num_dims must be in [1, 4] and perm must be a bijection over [0, num_dims - 1].",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_transpose_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions. The first params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[perm[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_transpose_params_f32 *params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output\n)",
              "source": {
                "line": 1135,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1135"
              },
              "summary": "Transpose a floating-point tensor."
            },
            {
              "description": "Concatenate tensors along the X axis.\n\nCall once per input tensor: `offset_x` selects where the input is stored along the X axis of the output tensor and must be advanced by `input_x` after each call. The output tensor must have the same height, channels and batch size as every input tensor.",
              "examples": [],
              "id": "arm_concatenation_f32_x",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f32_x",
              "params": [
                {
                  "description": "Pointer to the input tensor. Must not overlap the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Width of the input tensor.",
                  "direction": "in",
                  "name": "input_x",
                  "type": "int32_t"
                },
                {
                  "description": "Height of the input tensor.",
                  "direction": "in",
                  "name": "input_y",
                  "type": "int32_t"
                },
                {
                  "description": "Channels in the input tensor.",
                  "direction": "in",
                  "name": "input_z",
                  "type": "int32_t"
                },
                {
                  "description": "Batch size in the input tensor.",
                  "direction": "in",
                  "name": "input_w",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Width of the output tensor.",
                  "direction": "in",
                  "name": "output_x",
                  "type": "int32_t"
                },
                {
                  "description": "Offset on the X axis at which the input tensor is stored. Must be less than `output_x`.",
                  "direction": "in",
                  "name": "offset_x",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_f32_x(\n    const float32_t *input,\n    int32_t input_x,\n    int32_t input_y,\n    int32_t input_z,\n    int32_t input_w,\n    float32_t *output,\n    int32_t output_x,\n    uint32_t offset_x\n)",
              "source": {
                "line": 1158,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1158"
              },
              "summary": "Concatenate tensors along the X axis."
            },
            {
              "description": "Concatenate tensors along the Y axis.\n\nCall once per input tensor: `offset_y` selects where the input is stored along the Y axis of the output tensor and must be advanced by `input_y` after each call. The output tensor must have the same width, channels and batch size as every input tensor.",
              "examples": [],
              "id": "arm_concatenation_f32_y",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f32_y",
              "params": [
                {
                  "description": "Pointer to the input tensor. Must not overlap the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Width of the input tensor.",
                  "direction": "in",
                  "name": "input_x",
                  "type": "int32_t"
                },
                {
                  "description": "Height of the input tensor.",
                  "direction": "in",
                  "name": "input_y",
                  "type": "int32_t"
                },
                {
                  "description": "Channels in the input tensor.",
                  "direction": "in",
                  "name": "input_z",
                  "type": "int32_t"
                },
                {
                  "description": "Batch size in the input tensor.",
                  "direction": "in",
                  "name": "input_w",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Height of the output tensor.",
                  "direction": "in",
                  "name": "output_y",
                  "type": "int32_t"
                },
                {
                  "description": "Offset on the Y axis at which the input tensor is stored. Must be less than `output_y`.",
                  "direction": "in",
                  "name": "offset_y",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_f32_y(\n    const float32_t *input,\n    int32_t input_x,\n    int32_t input_y,\n    int32_t input_z,\n    int32_t input_w,\n    float32_t *output,\n    int32_t output_y,\n    uint32_t offset_y\n)",
              "source": {
                "line": 1183,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1183"
              },
              "summary": "Concatenate tensors along the Y axis."
            },
            {
              "description": "Concatenate tensors along the Z axis.\n\nCall once per input tensor: `offset_z` selects where the input is stored along the Z axis of the output tensor and must be advanced by `input_z` after each call. The output tensor must have the same width, height and batch size as every input tensor.",
              "examples": [],
              "id": "arm_concatenation_f32_z",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f32_z",
              "params": [
                {
                  "description": "Pointer to the input tensor. Must not overlap the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Width of the input tensor.",
                  "direction": "in",
                  "name": "input_x",
                  "type": "int32_t"
                },
                {
                  "description": "Height of the input tensor.",
                  "direction": "in",
                  "name": "input_y",
                  "type": "int32_t"
                },
                {
                  "description": "Channels in the input tensor.",
                  "direction": "in",
                  "name": "input_z",
                  "type": "int32_t"
                },
                {
                  "description": "Batch size in the input tensor.",
                  "direction": "in",
                  "name": "input_w",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Channels in the output tensor.",
                  "direction": "in",
                  "name": "output_z",
                  "type": "int32_t"
                },
                {
                  "description": "Offset on the Z axis at which the input tensor is stored. Must be less than `output_z`.",
                  "direction": "in",
                  "name": "offset_z",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_f32_z(\n    const float32_t *input,\n    int32_t input_x,\n    int32_t input_y,\n    int32_t input_z,\n    int32_t input_w,\n    float32_t *output,\n    int32_t output_z,\n    uint32_t offset_z\n)",
              "source": {
                "line": 1208,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1208"
              },
              "summary": "Concatenate tensors along the Z axis."
            },
            {
              "description": "Concatenate tensors along the W axis.\n\nCall once per input tensor: `offset_w` selects where the input is stored along the W axis of the output tensor and must be advanced by `input_w` after each call. The output tensor must have the same width, height and channels as every input tensor.",
              "examples": [],
              "id": "arm_concatenation_f32_w",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f32_w",
              "params": [
                {
                  "description": "Pointer to the input tensor. Must not overlap the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Width of the input tensor.",
                  "direction": "in",
                  "name": "input_x",
                  "type": "int32_t"
                },
                {
                  "description": "Height of the input tensor.",
                  "direction": "in",
                  "name": "input_y",
                  "type": "int32_t"
                },
                {
                  "description": "Channels in the input tensor.",
                  "direction": "in",
                  "name": "input_z",
                  "type": "int32_t"
                },
                {
                  "description": "Batch size in the input tensor.",
                  "direction": "in",
                  "name": "input_w",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Offset on the W axis at which the input tensor is stored.",
                  "direction": "in",
                  "name": "offset_w",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_f32_w(\n    const float32_t *input,\n    int32_t input_x,\n    int32_t input_y,\n    int32_t input_z,\n    int32_t input_w,\n    float32_t *output,\n    uint32_t offset_w\n)",
              "source": {
                "line": 1232,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1232"
              },
              "summary": "Concatenate tensors along the W axis."
            },
            {
              "description": "Concatenate float32 tensors of any rank along one axis.\n\nRank-agnostic sibling of the 4-D per-axis arm_concatenation_f32_{x,y,z,w} entry points: all inputs at once, any rank, any axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Input `s` has the output shape with `output_shape`[axis] replaced by `axis_sizes`[s]; the inputs are laid down in order along the axis. Inputs must not overlap the output. A dimension of 0 is accepted and copies nothing.",
              "examples": [],
              "id": "arm_concatenation_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f32",
              "params": [
                {
                  "description": "Array of `num_inputs` pointers to the flattened (row-major) inputs.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *const *"
                },
                {
                  "description": "Number of inputs (>= 1).",
                  "direction": "in",
                  "name": "num_inputs",
                  "type": "int32_t"
                },
                {
                  "description": "Array of length `num_inputs:` each input's extent along `axis`.",
                  "direction": "in",
                  "name": "axis_sizes",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of dimensions in `output_shape` (>= 1).",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "int32_t"
                },
                {
                  "description": "Output shape; `output_shape`[axis] must equal the sum of `axis_sizes`.",
                  "direction": "in",
                  "name": "output_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Axis to concatenate along (0 <= axis < output_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the flattened output.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, size entry, size sum, NULL pointer or an element count above INT32_MAX."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_concatenation_f32(\n    const float32_t *const *input_data,\n    int32_t num_inputs,\n    const int32_t *axis_sizes,\n    int32_t output_dims,\n    const int32_t *output_shape,\n    int32_t axis,\n    float32_t *output_data\n)",
              "source": {
                "line": 1259,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1259"
              },
              "summary": "Concatenate float32 tensors of any rank along one axis."
            },
            {
              "description": "Split a float32 tensor of any rank into several tensors along one axis.\n\nInverse of arm_concatenation_f32; per-split lengths also cover SPLIT_V. Output `s` has the input shape with `input_shape`[axis] replaced by `split_dims`[s]. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input. A dimension of 0 is accepted and copies nothing.",
              "examples": [],
              "id": "arm_split_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_split_f32",
              "params": [
                {
                  "description": "Pointer to the flattened (row-major) input.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of dimensions in `input_shape` (>= 1).",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "int32_t"
                },
                {
                  "description": "Input shape; `input_shape`[axis] must equal the sum of `split_dims`.",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Axis to split along (0 <= axis < input_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "Number of outputs (>= 1).",
                  "direction": "in",
                  "name": "num_splits",
                  "type": "int32_t"
                },
                {
                  "description": "Array of length `num_splits:` each output's extent along `axis`.",
                  "direction": "in",
                  "name": "split_dims",
                  "type": "const int32_t *"
                },
                {
                  "description": "Array of `num_splits` pointers to the flattened outputs.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *const *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, split entry, split sum, NULL pointer or an element count above INT32_MAX."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_split_f32(\n    const float32_t *input_data,\n    int32_t input_dims,\n    const int32_t *input_shape,\n    int32_t axis,\n    int32_t num_splits,\n    const int32_t *split_dims,\n    float32_t *const *output_data\n)",
              "source": {
                "line": 1285,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1285"
              },
              "summary": "Split a float32 tensor of any rank into several tensors along one axis."
            },
            {
              "description": "Stack float32 tensors of equal shape along a new axis (TFLite PACK).\n\nThe output shape is `input_shape` with `num_inputs` inserted at `axis`; input `s` lands at index `s` of that axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Inputs must not overlap the output. Rank-0 inputs (`input_dims` == 0, `axis` == 0) stack into a vector.",
              "examples": [],
              "id": "arm_pack_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_pack_f32",
              "params": [
                {
                  "description": "Array of `num_inputs` pointers to the flattened (row-major) inputs.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *const *"
                },
                {
                  "description": "Number of inputs (>= 1).",
                  "direction": "in",
                  "name": "num_inputs",
                  "type": "int32_t"
                },
                {
                  "description": "Number of dimensions of each input (>= 0).",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "int32_t"
                },
                {
                  "description": "Shape shared by every input (may be NULL when `input_dims` is 0).",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Position of the new axis in the output (0 <= axis <= input_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the flattened output.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, NULL pointer or an element count above INT32_MAX."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_pack_f32(\n    const float32_t *const *input_data,\n    int32_t num_inputs,\n    int32_t input_dims,\n    const int32_t *input_shape,\n    int32_t axis,\n    float32_t *output_data\n)",
              "source": {
                "line": 1310,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1310"
              },
              "summary": "Stack float32 tensors of equal shape along a new axis (TFLite PACK)."
            },
            {
              "description": "Unstack a float32 tensor along one axis into `input_shape`[axis] tensors (TFLite UNPACK).\n\nInverse of arm_pack_f32: output `s` is the input with the axis fixed at index `s` and removed from the shape. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input.",
              "examples": [],
              "id": "arm_unpack_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_unpack_f32",
              "params": [
                {
                  "description": "Pointer to the flattened (row-major) input.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of dimensions in `input_shape` (>= 1).",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "int32_t"
                },
                {
                  "description": "Input shape; `input_shape`[axis] (>= 1) is the number of outputs.",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Axis to unstack (0 <= axis < input_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "Array of `input_shape`[axis] pointers to the flattened outputs.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *const *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, a zero-extent unstack axis (no outputs to produce), NULL pointer or an element count above INT32_MAX."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_unpack_f32(\n    const float32_t *input_data,\n    int32_t input_dims,\n    const int32_t *input_shape,\n    int32_t axis,\n    float32_t *const *output_data\n)",
              "source": {
                "line": 1334,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1334"
              },
              "summary": "Unstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK)."
            },
            {
              "description": "Apply batch normalization.\n\nComputes `output = input * scale[c] + bias[c]` for every element of channel `c`, with `scale` and `bias` holding the pre-folded per-channel factors.",
              "examples": [],
              "id": "arm_batch_norm_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_norm_f32",
              "params": [
                {
                  "description": "Pointer to the input tensor data. Format: [N, H, W, C].",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output tensor data, same shape as `input`.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Per-channel scale, `input_dims->c` values.",
                  "direction": "in",
                  "name": "scale",
                  "type": "const float32_t *"
                },
                {
                  "description": "Per-channel bias, `input_dims->c` values.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float32_t *"
                },
                {
                  "description": "Input tensor dimensions. Every dimension must be positive.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector. Must be `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_batch_norm_f32(\n    const float32_t *input,\n    float32_t *output,\n    const float32_t *scale,\n    const float32_t *bias,\n    const cmsis_nn_dims *input_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 1390,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1390"
              },
              "summary": "Apply batch normalization."
            },
            {
              "description": "Reshape by copying data without changing element order.",
              "examples": [],
              "id": "arm_reshape_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reshape_f32",
              "params": [
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output tensor data. Nothing is copied when it aliases `input`.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of elements to copy.",
                  "direction": "in",
                  "name": "total_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_reshape_f32(const float32_t *input, float32_t *output, uint32_t total_size)",
              "source": {
                "line": 1404,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1404"
              },
              "summary": "Reshape by copying data without changing element order."
            },
            {
              "description": "Transpose a floating-point tensor.",
              "examples": [],
              "id": "arm_transpose_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Transpose parameters, including permutation and layout information. num_dims must be in [1, 4] and perm must be a bijection over [0, num_dims - 1].",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_transpose_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions. The first params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[perm[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_transpose_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
              "source": {
                "line": 3161,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3161"
              },
              "summary": "Transpose a floating-point tensor."
            },
            {
              "description": "Concatenate tensors along the X axis.\n\nCall once per input tensor: `offset_x` selects where the input is stored along the X axis of the output tensor and must be advanced by `input_x` after each call. The output tensor must have the same height, channels and batch size as every input tensor.",
              "examples": [],
              "id": "arm_concatenation_f16_x",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f16_x",
              "params": [
                {
                  "description": "Pointer to the input tensor. Must not overlap the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Width of the input tensor.",
                  "direction": "in",
                  "name": "input_x",
                  "type": "int32_t"
                },
                {
                  "description": "Height of the input tensor.",
                  "direction": "in",
                  "name": "input_y",
                  "type": "int32_t"
                },
                {
                  "description": "Channels in the input tensor.",
                  "direction": "in",
                  "name": "input_z",
                  "type": "int32_t"
                },
                {
                  "description": "Batch size in the input tensor.",
                  "direction": "in",
                  "name": "input_w",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Width of the output tensor.",
                  "direction": "in",
                  "name": "output_x",
                  "type": "int32_t"
                },
                {
                  "description": "Offset on the X axis at which the input tensor is stored. Must be less than `output_x`.",
                  "direction": "in",
                  "name": "offset_x",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_f16_x(\n    const float16_t *input,\n    int32_t input_x,\n    int32_t input_y,\n    int32_t input_z,\n    int32_t input_w,\n    float16_t *output,\n    int32_t output_x,\n    uint32_t offset_x\n)",
              "source": {
                "line": 3171,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3171"
              },
              "summary": "Concatenate tensors along the X axis."
            },
            {
              "description": "Concatenate tensors along the Y axis.\n\nCall once per input tensor: `offset_y` selects where the input is stored along the Y axis of the output tensor and must be advanced by `input_y` after each call. The output tensor must have the same width, channels and batch size as every input tensor.",
              "examples": [],
              "id": "arm_concatenation_f16_y",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f16_y",
              "params": [
                {
                  "description": "Pointer to the input tensor. Must not overlap the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Width of the input tensor.",
                  "direction": "in",
                  "name": "input_x",
                  "type": "int32_t"
                },
                {
                  "description": "Height of the input tensor.",
                  "direction": "in",
                  "name": "input_y",
                  "type": "int32_t"
                },
                {
                  "description": "Channels in the input tensor.",
                  "direction": "in",
                  "name": "input_z",
                  "type": "int32_t"
                },
                {
                  "description": "Batch size in the input tensor.",
                  "direction": "in",
                  "name": "input_w",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Height of the output tensor.",
                  "direction": "in",
                  "name": "output_y",
                  "type": "int32_t"
                },
                {
                  "description": "Offset on the Y axis at which the input tensor is stored. Must be less than `output_y`.",
                  "direction": "in",
                  "name": "offset_y",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_f16_y(\n    const float16_t *input,\n    int32_t input_x,\n    int32_t input_y,\n    int32_t input_z,\n    int32_t input_w,\n    float16_t *output,\n    int32_t output_y,\n    uint32_t offset_y\n)",
              "source": {
                "line": 3183,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3183"
              },
              "summary": "Concatenate tensors along the Y axis."
            },
            {
              "description": "Concatenate tensors along the Z axis.\n\nCall once per input tensor: `offset_z` selects where the input is stored along the Z axis of the output tensor and must be advanced by `input_z` after each call. The output tensor must have the same width, height and batch size as every input tensor.",
              "examples": [],
              "id": "arm_concatenation_f16_z",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f16_z",
              "params": [
                {
                  "description": "Pointer to the input tensor. Must not overlap the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Width of the input tensor.",
                  "direction": "in",
                  "name": "input_x",
                  "type": "int32_t"
                },
                {
                  "description": "Height of the input tensor.",
                  "direction": "in",
                  "name": "input_y",
                  "type": "int32_t"
                },
                {
                  "description": "Channels in the input tensor.",
                  "direction": "in",
                  "name": "input_z",
                  "type": "int32_t"
                },
                {
                  "description": "Batch size in the input tensor.",
                  "direction": "in",
                  "name": "input_w",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Channels in the output tensor.",
                  "direction": "in",
                  "name": "output_z",
                  "type": "int32_t"
                },
                {
                  "description": "Offset on the Z axis at which the input tensor is stored. Must be less than `output_z`.",
                  "direction": "in",
                  "name": "offset_z",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_f16_z(\n    const float16_t *input,\n    int32_t input_x,\n    int32_t input_y,\n    int32_t input_z,\n    int32_t input_w,\n    float16_t *output,\n    int32_t output_z,\n    uint32_t offset_z\n)",
              "source": {
                "line": 3195,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3195"
              },
              "summary": "Concatenate tensors along the Z axis."
            },
            {
              "description": "Concatenate tensors along the W axis.\n\nCall once per input tensor: `offset_w` selects where the input is stored along the W axis of the output tensor and must be advanced by `input_w` after each call. The output tensor must have the same width, height and channels as every input tensor.",
              "examples": [],
              "id": "arm_concatenation_f16_w",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f16_w",
              "params": [
                {
                  "description": "Pointer to the input tensor. Must not overlap the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Width of the input tensor.",
                  "direction": "in",
                  "name": "input_x",
                  "type": "int32_t"
                },
                {
                  "description": "Height of the input tensor.",
                  "direction": "in",
                  "name": "input_y",
                  "type": "int32_t"
                },
                {
                  "description": "Channels in the input tensor.",
                  "direction": "in",
                  "name": "input_z",
                  "type": "int32_t"
                },
                {
                  "description": "Batch size in the input tensor.",
                  "direction": "in",
                  "name": "input_w",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output tensor.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Offset on the W axis at which the input tensor is stored.",
                  "direction": "in",
                  "name": "offset_w",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_f16_w(\n    const float16_t *input,\n    int32_t input_x,\n    int32_t input_y,\n    int32_t input_z,\n    int32_t input_w,\n    float16_t *output,\n    uint32_t offset_w\n)",
              "source": {
                "line": 3207,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3207"
              },
              "summary": "Concatenate tensors along the W axis."
            },
            {
              "description": "Concatenate float32 tensors of any rank along one axis.\n\nRank-agnostic sibling of the 4-D per-axis arm_concatenation_f32_{x,y,z,w} entry points: all inputs at once, any rank, any axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Input `s` has the output shape with `output_shape`[axis] replaced by `axis_sizes`[s]; the inputs are laid down in order along the axis. Inputs must not overlap the output. A dimension of 0 is accepted and copies nothing.",
              "examples": [],
              "id": "arm_concatenation_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_f16",
              "params": [
                {
                  "description": "Array of `num_inputs` pointers to the flattened (row-major) inputs.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *const *"
                },
                {
                  "description": "Number of inputs (>= 1).",
                  "direction": "in",
                  "name": "num_inputs",
                  "type": "int32_t"
                },
                {
                  "description": "Array of length `num_inputs:` each input's extent along `axis`.",
                  "direction": "in",
                  "name": "axis_sizes",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of dimensions in `output_shape` (>= 1).",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "int32_t"
                },
                {
                  "description": "Output shape; `output_shape`[axis] must equal the sum of `axis_sizes`.",
                  "direction": "in",
                  "name": "output_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Axis to concatenate along (0 <= axis < output_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the flattened output.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, size entry, size sum, NULL pointer or an element count above INT32_MAX."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_concatenation_f16(\n    const float16_t *const *input_data,\n    int32_t num_inputs,\n    const int32_t *axis_sizes,\n    int32_t output_dims,\n    const int32_t *output_shape,\n    int32_t axis,\n    float16_t *output_data\n)",
              "source": {
                "line": 3218,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3218"
              },
              "summary": "Concatenate float32 tensors of any rank along one axis."
            },
            {
              "description": "Stack float32 tensors of equal shape along a new axis (TFLite PACK).\n\nThe output shape is `input_shape` with `num_inputs` inserted at `axis`; input `s` lands at index `s` of that axis. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Inputs must not overlap the output. Rank-0 inputs (`input_dims` == 0, `axis` == 0) stack into a vector.",
              "examples": [],
              "id": "arm_pack_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_pack_f16",
              "params": [
                {
                  "description": "Array of `num_inputs` pointers to the flattened (row-major) inputs.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *const *"
                },
                {
                  "description": "Number of inputs (>= 1).",
                  "direction": "in",
                  "name": "num_inputs",
                  "type": "int32_t"
                },
                {
                  "description": "Number of dimensions of each input (>= 0).",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "int32_t"
                },
                {
                  "description": "Shape shared by every input (may be NULL when `input_dims` is 0).",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Position of the new axis in the output (0 <= axis <= input_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the flattened output.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (output untouched) on an invalid rank, axis, shape entry, NULL pointer or an element count above INT32_MAX."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_pack_f16(\n    const float16_t *const *input_data,\n    int32_t num_inputs,\n    int32_t input_dims,\n    const int32_t *input_shape,\n    int32_t axis,\n    float16_t *output_data\n)",
              "source": {
                "line": 3229,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3229"
              },
              "summary": "Stack float32 tensors of equal shape along a new axis (TFLite PACK)."
            },
            {
              "description": "Unstack a float32 tensor along one axis into `input_shape`[axis] tensors (TFLite UNPACK).\n\nInverse of arm_pack_f32: output `s` is the input with the axis fixed at index `s` and removed from the shape. Bit copy, NaN/Inf/-0/subnormal payloads preserved. Outputs must not overlap the input.",
              "examples": [],
              "id": "arm_unpack_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_unpack_f16",
              "params": [
                {
                  "description": "Pointer to the flattened (row-major) input.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of dimensions in `input_shape` (>= 1).",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "int32_t"
                },
                {
                  "description": "Input shape; `input_shape`[axis] (>= 1) is the number of outputs.",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Axis to unstack (0 <= axis < input_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "Array of `input_shape`[axis] pointers to the flattened outputs.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *const *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` (outputs untouched) on an invalid rank, axis, shape entry, a zero-extent unstack axis (no outputs to produce), NULL pointer or an element count above INT32_MAX."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_unpack_f16(\n    const float16_t *input_data,\n    int32_t input_dims,\n    const int32_t *input_shape,\n    int32_t axis,\n    float16_t *const *output_data\n)",
              "source": {
                "line": 3239,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3239"
              },
              "summary": "Unstack a float32 tensor along one axis into inputshape[axis] tensors (TFLite UNPACK)."
            },
            {
              "description": "Apply batch normalization.\n\nComputes `output = input * scale[c] + bias[c]` for every element of channel `c`, with `scale` and `bias` holding the pre-folded per-channel factors.",
              "examples": [],
              "id": "arm_batch_norm_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_norm_f16",
              "params": [
                {
                  "description": "Pointer to the input tensor data. Format: [N, H, W, C].",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output tensor data, same shape as `input`.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Per-channel scale, `input_dims->c` values.",
                  "direction": "in",
                  "name": "scale",
                  "type": "const float16_t *"
                },
                {
                  "description": "Per-channel bias, `input_dims->c` values.",
                  "direction": "in",
                  "name": "bias",
                  "type": "const float16_t *"
                },
                {
                  "description": "Input tensor dimensions. Every dimension must be positive.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Tensor layout selector. Must be `ARM_NN_LAYOUT_NHWC`.",
                  "direction": "in",
                  "name": "layout",
                  "type": "arm_nn_tensor_layout"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_batch_norm_f16(\n    const float16_t *input,\n    float16_t *output,\n    const float16_t *scale,\n    const float16_t *bias,\n    const cmsis_nn_dims *input_dims,\n    arm_nn_tensor_layout layout\n)",
              "source": {
                "line": 3272,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3272"
              },
              "summary": "Apply batch normalization."
            },
            {
              "description": "Reshape by copying data without changing element order.",
              "examples": [],
              "id": "arm_reshape_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reshape_f16",
              "params": [
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output tensor data. Nothing is copied when it aliases `input`.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Number of elements to copy.",
                  "direction": "in",
                  "name": "total_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_reshape_f16(const float16_t *input, float16_t *output, uint32_t total_size)",
              "source": {
                "line": 3282,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3282"
              },
              "summary": "Reshape by copying data without changing element order."
            }
          ]
        },
        {
          "description": "",
          "name": "Pad Layer Functions:",
          "path": "heliaCORE.Pad",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Pad a tensor with a constant value.",
              "examples": [],
              "id": "arm_pad_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_pad_f32",
              "params": [
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to the output tensor data, sized by `input_size` plus `pre_pad` and `post_pad` in every dimension.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Value to pad with.",
                  "direction": "in",
                  "name": "pad_value",
                  "type": "float32_t"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_size",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Padding to apply before the data in each dimension.",
                  "direction": "in",
                  "name": "pre_pad",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Padding to apply after the data in each dimension.",
                  "direction": "in",
                  "name": "post_pad",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL or a padded output dimension is not positive."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_pad_f32(\n    const float32_t *input,\n    float32_t *output,\n    float32_t pad_value,\n    const cmsis_nn_dims *input_size,\n    const cmsis_nn_dims *pre_pad,\n    const cmsis_nn_dims *post_pad\n)",
              "source": {
                "line": 1361,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1361"
              },
              "summary": "Pad a tensor with a constant value."
            },
            {
              "description": "Pad a tensor with a constant value.",
              "examples": [],
              "id": "arm_pad_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_pad_f16",
              "params": [
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the output tensor data, sized by `input_size` plus `pre_pad` and `post_pad` in every dimension.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                },
                {
                  "description": "Value to pad with.",
                  "direction": "in",
                  "name": "pad_value",
                  "type": "float16_t"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_size",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Padding to apply before the data in each dimension.",
                  "direction": "in",
                  "name": "pre_pad",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Padding to apply after the data in each dimension.",
                  "direction": "in",
                  "name": "post_pad",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success, or `ARM_CMSIS_NN_ARG_ERROR` when a pointer is NULL or a padded output dimension is not positive."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_pad_f16(\n    const float16_t *input,\n    float16_t *output,\n    float16_t pad_value,\n    const cmsis_nn_dims *input_size,\n    const cmsis_nn_dims *pre_pad,\n    const cmsis_nn_dims *post_pad\n)",
              "source": {
                "line": 3255,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3255"
              },
              "summary": "Pad a tensor with a constant value."
            }
          ]
        },
        {
          "description": "Perform max and average pooling operations",
          "name": "Pooling Functions",
          "path": "heliaCORE.Pooling",
          "submodules": [],
          "summary": "Perform max and average pooling operations",
          "symbols": [
            {
              "description": "Max pooling.",
              "examples": [],
              "id": "arm_max_pool_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_max_pool_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pooling parameters (stride, padding and activation clamp).",
                  "direction": "in",
                  "name": "pool_params",
                  "type": "const cmsis_nn_pool_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "src",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pooling kernel dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success, including an output with no rows or no columns (an extent of 0 or less), which writes nothing; `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments: a NULL pointer argument other than ctx, a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_max_pool_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_pool_params_f32 *pool_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *src,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    float32_t *dst\n)",
              "source": {
                "line": 540,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L540"
              },
              "summary": "Max pooling."
            },
            {
              "description": "Average pooling.",
              "examples": [],
              "id": "arm_avg_pool_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avg_pool_f32",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pooling parameters (stride, padding and activation clamp).",
                  "direction": "in",
                  "name": "pool_params",
                  "type": "const cmsis_nn_pool_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "src",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pooling kernel dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success, including an output with no rows or no columns (an extent of 0 or less), which writes nothing; `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments: a NULL pointer argument other than ctx, a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_avg_pool_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_pool_params_f32 *pool_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *src,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    float32_t *dst\n)",
              "source": {
                "line": 565,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L565"
              },
              "summary": "Average pooling."
            },
            {
              "description": "Max pooling.\n\n:::note\nThe output activation clamp on the scalar (non-MVE) build path is the bit-classified clamp of #380, so a NaN that reaches the clamp comes back as NaN at every optimization level on the gated toolchains rather than as a bound. A NaN rarely reaches it, though: the scalar max reduction uses an ordered compare that drops a NaN window element (and its NaN behavior at the shipped -Ofast is unspecified), and the MVE path's vmaxnmq reduction and vmaxnmq/vminnmq clamp suppress NaN, so this kernel does not promise NaN propagation end to end.\n\n:::",
              "examples": [],
              "id": "arm_max_pool_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_max_pool_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pooling parameters (stride, padding and activation clamp).",
                  "direction": "in",
                  "name": "pool_params",
                  "type": "const cmsis_nn_pool_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "src",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pooling kernel dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success, including an output with no rows or no columns (an extent of 0 or less), which writes nothing; `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments: a NULL pointer argument other than ctx, a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_max_pool_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_pool_params_f16 *pool_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *src,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    float16_t *dst\n)",
              "source": {
                "line": 2695,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2695"
              },
              "summary": "Max pooling."
            },
            {
              "description": "Average pooling.\n\n:::note\nOn non-MVE builds every output element goes through the bit-classified scalar clamp of #380, so a NaN in the pooling window propagates through the window sum and the output activation clamp to the output element at every optimization level on the gated toolchains, including the shipped -Ofast. On MVE builds the clamp is vmaxnmq/vminnmq with no NaN restore, so a NaN resolves to a clamp bound there instead.\n\n:::",
              "examples": [],
              "id": "arm_avg_pool_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avg_pool_f16",
              "params": [
                {
                  "description": "Function context that may hold a temporary scratch buffer.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pooling parameters (stride, padding and activation clamp).",
                  "direction": "in",
                  "name": "pool_params",
                  "type": "const cmsis_nn_pool_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "src",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pooling kernel dimensions.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "dst",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success, including an output with no rows or no columns (an extent of 0 or less), which writes nothing; `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments: a NULL pointer argument other than ctx, a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_avg_pool_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_pool_params_f16 *pool_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *src,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    float16_t *dst\n)",
              "source": {
                "line": 2712,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2712"
              },
              "summary": "Average pooling."
            }
          ]
        },
        {
          "description": "A collection of functions to perform basic operations for neural network layers. Functions with a _s8 suffix support TensorFlow Lite framework.",
          "name": "Public",
          "path": "heliaCORE.Public",
          "submodules": [],
          "summary": "A collection of functions to perform basic operations for neural network layers.",
          "symbols": []
        },
        {
          "description": "",
          "name": "Quantization Functions:",
          "path": "heliaCORE.Quantization",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Widen a float16 vector to float32.\n\nBit-exact widening of every input class: finite values, subnormals (normal in float32), +/-0 and +/-Inf convert exactly. No accumulation, no rounding. NaN behavior: on every leg a NaN stays a NaN with its sign, quiet bit and payload preserved bit-exactly (a signaling NaN stays signaling). The scalar leg widens on integer lanes and raises no floating-point exception flag. The MVE leg converts each 8-element block with the vector VCVT first and then rebuilds the NaN lanes from the half's bits (per 4-lane vector, 8 elements per main-loop block), so a signaling-NaN input may leave FPSCR.IOC (invalid operation, cumulative) set on that leg; no trap, and the result is the same bits. Input and output must not overlap. Serves the f16-weights DEQUANTIZE op (`kws_float_fp16_weights`).",
              "examples": [],
              "id": "arm_dequantize_f16_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_dequantize_f16_f32",
              "params": [
                {
                  "description": "Pointer to the float16 input vector.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Pointer to the float32 output vector.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                },
                {
                  "description": "Number of elements (0 is a no-op).",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS`, or `ARM_CMSIS_NN_ARG_ERROR` when `block_size` is negative or a pointer is NULL with a non-zero `block_size`."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_dequantize_f16_f32(const float16_t *input, float32_t *output, int32_t block_size)",
              "source": {
                "line": 2918,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2918"
              },
              "summary": "Widen a float16 vector to float32."
            }
          ]
        },
        {
          "description": "",
          "name": "Reduction Functions",
          "path": "heliaCORE.Reduction",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Computes the sum of the input tensor along the specified axes.\n\nSums are accumulated in float32 (also for the float16 variant, which rounds once to float16 at the end), so results do not overflow at float16 range and precision does not degrade with the reduction count. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs.",
              "examples": [],
              "id": "arm_reduce_sum_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_sum_f32",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Input tensor dimensions (4D NHWC)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "4D binary axis mask (non-zero = reduce that axis)",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Output tensor dimensions (reduced axes have size 1)",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_sum_f32(\n    const float32_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1902,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1902"
              },
              "summary": "Computes the sum of the input tensor along the specified axes."
            },
            {
              "description": "Returns the first minimum's axis-relative INT32 index for a f32 tensor.\n\nThe input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.\n\nA NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.\n\nMetadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.",
              "examples": [],
              "id": "arm_argmin_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_argmin_f32",
              "params": [
                {
                  "description": "Input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Four NHWC extents.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Canonical reduction axis, in [0,3].",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "INT32 indices, each in [0,input_dims[axis]).",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_argmin_f32(\n    const float32_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    int32_t axis,\n    int32_t *output_data\n)",
              "source": {
                "line": 1940,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1940"
              },
              "summary": "Returns the first minimum's axis-relative INT32 index for a f32 tensor."
            },
            {
              "description": "Returns the first maximum's axis-relative INT32 index for a f32 tensor.\n\nThe input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.\n\nA NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.\n\nMetadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.",
              "examples": [],
              "id": "arm_argmax_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_argmax_f32",
              "params": [
                {
                  "description": "Input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Four NHWC extents.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Canonical reduction axis, in [0,3].",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "INT32 indices, each in [0,input_dims[axis]).",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_argmax_f32(\n    const float32_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    int32_t axis,\n    int32_t *output_data\n)",
              "source": {
                "line": 1974,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1974"
              },
              "summary": "Returns the first maximum's axis-relative INT32 index for a f32 tensor."
            },
            {
              "description": "Reduces a f32 NHWC tensor to its maximum along a binary axis mask.\n\nValues are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7fc00000); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.\n\nA zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces -Inf; an empty output performs no accesses to data buffers.\n\nAll metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.",
              "examples": [],
              "id": "arm_reduce_max_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_max_f32",
              "params": [
                {
                  "description": "Input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Four NHWC extents.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Four binary reduction flags.",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "NHWC output shape with reduced axes retained as 1.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_max_f32(\n    const float32_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2004,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2004"
              },
              "summary": "Reduces a f32 NHWC tensor to its maximum along a binary axis mask."
            },
            {
              "description": "Reduces a f32 NHWC tensor to its minimum along a binary axis mask.\n\nValues are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7fc00000); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.\n\nA zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces +Inf; an empty output performs no accesses to data buffers.\n\nAll metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.",
              "examples": [],
              "id": "arm_reduce_min_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_min_f32",
              "params": [
                {
                  "description": "Input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Four NHWC extents.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Four binary reduction flags.",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "NHWC output shape with reduced axes retained as 1.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_min_f32(\n    const float32_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2038,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2038"
              },
              "summary": "Reduces a f32 NHWC tensor to its minimum along a binary axis mask."
            },
            {
              "description": "Computes the mean of a float32 tensor along the specified axes.\n\nValues are accumulated and divided once in float32; unlike the float16 variant there is no wider accumulator, so rounding error can grow with the reduction length, matching arm_reduce_sum_f32. Because the intermediate accumulation is itself float32, it can saturate to +/-Inf even when the mean itself is representable, but whether it does depends on accumulation order: a strictly sequential build keeps one running sum, while vector builds  MVE intrinsics, or compiler auto-vectorization of the scalar path at -Ofast  fold per-lane partial sums, so on inputs whose partial sums exceed FLT_MAX in magnitude either build may return +/-Inf and the two may disagree (one finite, one Inf); when partial sums of opposite sign both saturate, the vector fold can even yield NaN (Inf + -Inf) from all-finite inputs. Only when every accumulation order overflows  e.g. same-signed values summing past FLT_MAX  is +/-Inf guaranteed on all builds. This is the accumulation-order divergence described below taken to the extreme. NaN and Inf propagate. A mean over all -0.0f inputs returns +0.0f on every build: the accumulator starts at +0.0f and (+0.0f) + (-0.0f) is +0.0f under round-to-nearest. Vector and scalar builds may differ in final ulps because float accumulation order differs.\n\nUnlike arm_reduce_sum_f32 (identical signature, null checks only), this kernel validates shapes and returns `ARM_CMSIS_NN_ARG_ERROR` when any input dimension is less than 1, when any `output_dims` entry differs from the input shape with the reduced axes collapsed to 1, or when the input element count or the reduction count does not fit in int32_t. `output_data` must not overlap `input_data:` each output element is written after reading its whole reduction set, so an aliased write can corrupt inputs still to be read.",
              "examples": [],
              "id": "arm_nn_mean_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mean_f32",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Input tensor dimensions (4D NHWC)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "4D binary axis mask (non-zero = reduce that axis)",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Output tensor dimensions (reduced axes have size 1)",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mean_f32(\n    const float32_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2086,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2086"
              },
              "summary": "Computes the mean of a float32 tensor along the specified axes."
            },
            {
              "description": "Computes the sum of the input tensor along the specified axes.\n\nSums are accumulated in float32 (also for the float16 variant, which rounds once to float16 at the end), so results do not overflow at float16 range and precision does not degrade with the reduction count. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs.",
              "examples": [],
              "id": "arm_reduce_sum_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_sum_f16",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Input tensor dimensions (4D NHWC)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "4D binary axis mask (non-zero = reduce that axis)",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Output tensor dimensions (reduced axes have size 1)",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_sum_f16(\n    const float16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 3646,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3646"
              },
              "summary": "Computes the sum of the input tensor along the specified axes."
            },
            {
              "description": "Returns the first minimum's axis-relative INT32 index for a f16 tensor.\n\nThe input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.\n\nA NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.\n\nMetadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.",
              "examples": [],
              "id": "arm_argmin_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_argmin_f16",
              "params": [
                {
                  "description": "Input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Four NHWC extents.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Canonical reduction axis, in [0,3].",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "INT32 indices, each in [0,input_dims[axis]).",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_argmin_f16(\n    const float16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    int32_t axis,\n    int32_t *output_data\n)",
              "source": {
                "line": 3684,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3684"
              },
              "summary": "Returns the first minimum's axis-relative INT32 index for a f16 tensor."
            },
            {
              "description": "Returns the first maximum's axis-relative INT32 index for a f16 tensor.\n\nThe input is contiguous NHWC with four extents; axis is a canonical index 0..3. Output contains the product of the other three extents, in row-major order with the reduced axis removed. Logical ranks, negative-axis normalization and squeezed output metadata are the caller's responsibility. No scratch is needed.\n\nA NaN never wins, regardless of payload, sign or signaling bit, as in LiteRT's reference ARG_MAX/ARG_MIN: the first non-NaN extremum is selected and an all-NaN line returns index 0. Equal numeric extrema retain the first index, including +0/-0 ties. Infinities and subnormals follow numeric order. Selection uses raw bits, with no floating-point arithmetic or conversion; numerical FP controls and cumulative exception flags are preserved. Native LiteRT FP16 evaluation is not implied.\n\nMetadata is required; extents must be nonnegative and the reduced extent must be positive, even when another extent is zero. Declared input and INT32 output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Valid empty outputs perform no data accesses. Buffers must be normally aligned, contiguous, adequately allocated and non-overlapping with each other and metadata; capacity and overlap are caller preconditions. All detected errors precede output writes.",
              "examples": [],
              "id": "arm_argmax_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_argmax_f16",
              "params": [
                {
                  "description": "Input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Four NHWC extents.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Canonical reduction axis, in [0,3].",
                  "direction": "in",
                  "name": "axis",
                  "type": "int32_t"
                },
                {
                  "description": "INT32 indices, each in [0,input_dims[axis]).",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_argmax_f16(\n    const float16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    int32_t axis,\n    int32_t *output_data\n)",
              "source": {
                "line": 3718,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3718"
              },
              "summary": "Returns the first maximum's axis-relative INT32 index for a f16 tensor."
            },
            {
              "description": "Reduces a f16 NHWC tensor to its maximum along a binary axis mask.\n\nValues are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7e00); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.\n\nA zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces -Inf; an empty output performs no accesses to data buffers.\n\nAll metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.",
              "examples": [],
              "id": "arm_reduce_max_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_max_f16",
              "params": [
                {
                  "description": "Input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Four NHWC extents.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Four binary reduction flags.",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "NHWC output shape with reduced axes retained as 1.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_max_f16(\n    const float16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 3748,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3748"
              },
              "summary": "Reduces a f16 NHWC tensor to its maximum along a binary axis mask."
            },
            {
              "description": "Reduces a f16 NHWC tensor to its minimum along a binary axis mask.\n\nValues are selected without floating-point arithmetic, accumulation or conversion. Any NaN in a reduction yields canonical quiet NaN (0x7e00); infinities and subnormals retain their bits. Equal numeric values retain the first input in row-major order, including zero signs. Scalar and MVE paths share this bit contract independently of FP controls. LiteRT nonfinite/zero-sign behavior may differ by shape/resolver. Refs #498.\n\nA zero mask copies bits unchanged, including NaN payloads. Reducing a singleton axis instead canonicalizes NaNs. An empty reduced domain produces +Inf; an empty output performs no accesses to data buffers.\n\nAll metadata pointers are required. Extents must be nonnegative, mask entries exactly 0 or 1, and output extents equal input extents with reduced axes retained as 1. Declared input/output byte counts must each fit INT32_MAX; any zero extent makes its tensor count zero. Data pointers may be NULL only for zero-element tensors. Buffers must be contiguous, normally aligned, adequately allocated and non-overlapping; allocation capacity and overlap are caller preconditions, not runtime checks.",
              "examples": [],
              "id": "arm_reduce_min_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_min_f16",
              "params": [
                {
                  "description": "Input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Four NHWC extents.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Four binary reduction flags.",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "NHWC output shape with reduced axes retained as 1.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR before any output write."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_min_f16(\n    const float16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 3782,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3782"
              },
              "summary": "Reduces a f16 NHWC tensor to its minimum along a binary axis mask."
            },
            {
              "description": "Computes the mean of a float16 tensor along the specified axes.\n\nValues are accumulated and divided in float32, then rounded once to float16. NaN and Inf propagate. Vector and scalar builds may differ in final ulps because float accumulation order differs. Builds at -Ofast (the shipped CMSIS_OPTIMIZATION_LEVEL) may additionally differ from lower optimization levels by 1 ulp for non-power-of-two reduction counts: -freciprocal-math turns the divide-by-count into a multiply-by-reciprocal, which rounds differently.\n\nUnlike arm_reduce_sum_f16 (identical signature, null checks only), this kernel validates shapes and returns `ARM_CMSIS_NN_ARG_ERROR` when any input dimension is less than 1, when any `output_dims` entry differs from the input shape with the reduced axes collapsed to 1, or when the input element count or the reduction count does not fit in int32_t. `output_data` must not overlap `input_data:` each output element is written after reading its whole reduction set, so an aliased write can corrupt inputs still to be read.",
              "examples": [],
              "id": "arm_nn_mean_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mean_f16",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Input tensor dimensions (4D NHWC)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "4D binary axis mask (non-zero = reduce that axis)",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Output tensor dimensions (reduced axes have size 1)",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mean_f16(\n    const float16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 3817,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3817"
              },
              "summary": "Computes the mean of a float16 tensor along the specified axes."
            }
          ]
        },
        {
          "description": "",
          "name": "Reshape Functions",
          "path": "heliaCORE.Reshape",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Scratch size in bytes for `arm_resize_nearest_neighbor_f32()` / `arm_resize_nearest_neighbor_f16()`.\n\nThe kernels precompute one int32_t input index per output row and per output column, so the requirement is (output_dims->h + output_dims->w) * sizeof(int32_t). Returns -1 (never 0) when `output_dims` is NULL, when h or w is less than 1, or when the size does not fit in int32_t; a negative result must not be used to size a buffer, and the kernels reject a { NULL, 0 } context outright (the -1 family of the integer sizers, not the 0-returning family most float sizers use; see the sentinel note on arm_nn_size_mul).",
              "examples": [],
              "id": "arm_resize_nearest_neighbor_f32_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_resize_nearest_neighbor_f32_get_buffer_size",
              "params": [
                {
                  "description": "Output tensor dimensions (only h and w are read).",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required ctx->size in bytes, or -1."
                }
              ],
              "signature": "int32_t arm_resize_nearest_neighbor_f32_get_buffer_size(const cmsis_nn_dims *output_dims)",
              "source": {
                "line": 1425,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1425"
              },
              "summary": "Scratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16()."
            },
            {
              "description": "Nearest-neighbor resize of a float32 NHWC tensor.\n\nPure data movement: every output element is a bit copy of one input element, so NaN (sign and payload), +/-Inf, -0.0 and subnormals are preserved bit-for-bit on the scalar and MVE legs alike (the MVE copy is a tail-predicated vldr/vstr pair, not an FP operation, so FPSCR flush-to-zero and default-NaN do not apply). No arithmetic is performed on the data; the only float math is the float32 index scale below.\n\nIndex semantics are TFLite's RESIZE_NEAREST_NEIGHBOR reference, evaluated in float32 per axis: scale = (align_corners && out > 1) ? (in - 1) / (out - 1) : in / out idx = align_corners ? roundf((o + offset) * scale) : floorf((o + offset) * scale) idx = min(idx, in - 1); if (half_pixel_centers) idx = max(idx, 0) with offset = half_pixel_centers ? 0.5f : 0.0f; roundf rounds ties away from zero, matching TfLiteRound. All four align_corners/half_pixel_centers combinations were verified bit-for-bit against TFLite 2.20 over a shape sweep (1..16 square, 80 random NHWC shapes up to 40x40, 224->7 and 7->224), including the out == 1 align_corners case, which maps to input index 0.",
              "examples": [],
              "id": "arm_resize_nearest_neighbor_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_resize_nearest_neighbor_f32",
              "params": [
                {
                  "description": "Scratch context. ctx->buf must be non-NULL and 4-byte aligned, and ctx->size at least arm_resize_nearest_neighbor_f32_get_buffer_size(output_shape); the kernel writes the x/y index maps here and does not read them after returning.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "align_corners / half_pixel_centers.",
                  "direction": "in",
                  "name": "resize_params",
                  "type": "const cmsis_nn_resize_params *"
                },
                {
                  "description": "Input tensor dimensions in NHWC format; every dimension must be >= 1.",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input tensor data. Must not overlap `output_data`.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Dimensions of the output-size tensor; must hold exactly 2 elements.",
                  "direction": "in",
                  "name": "output_size_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output size as [output_height, output_width], both >= 1.",
                  "direction": "in",
                  "name": "output_size_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions in NHWC format; n and c must equal the input's and h/w must equal `output_size_data`.",
                  "direction": "in",
                  "name": "output_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when any constraint above fails (including a NULL pointer argument); nothing is written on ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_resize_nearest_neighbor_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_resize_params *resize_params,\n    const cmsis_nn_dims *input_shape,\n    const float32_t *input_data,\n    const cmsis_nn_dims *output_size_shape,\n    const int32_t *output_size_data,\n    const cmsis_nn_dims *output_shape,\n    float32_t *output_data\n)",
              "source": {
                "line": 1458,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1458"
              },
              "summary": "Nearest-neighbor resize of a float32 NHWC tensor."
            },
            {
              "description": "Scratch size in bytes for `arm_resize_nearest_neighbor_f32()` / `arm_resize_nearest_neighbor_f16()`.\n\nThe kernels precompute one int32_t input index per output row and per output column, so the requirement is (output_dims->h + output_dims->w) * sizeof(int32_t). Returns -1 (never 0) when `output_dims` is NULL, when h or w is less than 1, or when the size does not fit in int32_t; a negative result must not be used to size a buffer, and the kernels reject a { NULL, 0 } context outright (the -1 family of the integer sizers, not the 0-returning family most float sizers use; see the sentinel note on arm_nn_size_mul).",
              "examples": [],
              "id": "arm_resize_nearest_neighbor_f16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_resize_nearest_neighbor_f16_get_buffer_size",
              "params": [
                {
                  "description": "Output tensor dimensions (only h and w are read).",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required ctx->size in bytes, or -1."
                }
              ],
              "signature": "int32_t arm_resize_nearest_neighbor_f16_get_buffer_size(const cmsis_nn_dims *output_dims)",
              "source": {
                "line": 3294,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3294"
              },
              "summary": "Scratch size in bytes for armresizenearestneighborf32() / armresizenearestneighborf16()."
            },
            {
              "description": "Nearest-neighbor resize of a float32 NHWC tensor.\n\nPure data movement: every output element is a bit copy of one input element, so NaN (sign and payload), +/-Inf, -0.0 and subnormals are preserved bit-for-bit on the scalar and MVE legs alike (the MVE copy is a tail-predicated vldr/vstr pair, not an FP operation, so FPSCR flush-to-zero and default-NaN do not apply). No arithmetic is performed on the data; the only float math is the float32 index scale below.\n\nIndex semantics are TFLite's RESIZE_NEAREST_NEIGHBOR reference, evaluated in float32 per axis: scale = (align_corners && out > 1) ? (in - 1) / (out - 1) : in / out idx = align_corners ? roundf((o + offset) * scale) : floorf((o + offset) * scale) idx = min(idx, in - 1); if (half_pixel_centers) idx = max(idx, 0) with offset = half_pixel_centers ? 0.5f : 0.0f; roundf rounds ties away from zero, matching TfLiteRound. All four align_corners/half_pixel_centers combinations were verified bit-for-bit against TFLite 2.20 over a shape sweep (1..16 square, 80 random NHWC shapes up to 40x40, 224->7 and 7->224), including the out == 1 align_corners case, which maps to input index 0.\n\n:::note\nfloat16 twin: each element is copied as a 16-bit lane with no widening or conversion, so half-precision NaN payloads and subnormals are preserved exactly and the data is never evaluated in float32. Scratch is sized by `arm_resize_nearest_neighbor_f16_get_buffer_size()` (same query as f32).\n\n:::",
              "examples": [],
              "id": "arm_resize_nearest_neighbor_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_resize_nearest_neighbor_f16",
              "params": [
                {
                  "description": "Scratch context. ctx->buf must be non-NULL and 4-byte aligned, and ctx->size at least arm_resize_nearest_neighbor_f32_get_buffer_size(output_shape); the kernel writes the x/y index maps here and does not read them after returning.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "align_corners / half_pixel_centers.",
                  "direction": "in",
                  "name": "resize_params",
                  "type": "const cmsis_nn_resize_params *"
                },
                {
                  "description": "Input tensor dimensions in NHWC format; every dimension must be >= 1.",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input tensor data. Must not overlap `output_data`.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Dimensions of the output-size tensor; must hold exactly 2 elements.",
                  "direction": "in",
                  "name": "output_size_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output size as [output_height, output_width], both >= 1.",
                  "direction": "in",
                  "name": "output_size_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions in NHWC format; n and c must equal the input's and h/w must equal `output_size_data`.",
                  "direction": "in",
                  "name": "output_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when any constraint above fails (including a NULL pointer argument); nothing is written on ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_resize_nearest_neighbor_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_resize_params *resize_params,\n    const cmsis_nn_dims *input_shape,\n    const float16_t *input_data,\n    const cmsis_nn_dims *output_size_shape,\n    const int32_t *output_size_data,\n    const cmsis_nn_dims *output_shape,\n    float16_t *output_data\n)",
              "source": {
                "line": 3302,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3302"
              },
              "summary": "Nearest-neighbor resize of a float32 NHWC tensor."
            }
          ]
        },
        {
          "description": "",
          "name": "Softmax Functions",
          "path": "heliaCORE.Softmax",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Softmax using the float-native API signature.",
              "examples": [],
              "id": "arm_softmax_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_softmax_f32",
              "params": [
                {
                  "description": "Pointer to the input matrix stored as `num_rows` rows of `row_size` values.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float32_t *"
                },
                {
                  "description": "Number of rows in the input matrix.",
                  "direction": "in",
                  "name": "num_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of columns per row.",
                  "direction": "in",
                  "name": "row_size",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output matrix.",
                  "direction": "out",
                  "name": "output",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_softmax_f32(const float32_t *input, int32_t num_rows, int32_t row_size, float32_t *output)",
              "source": {
                "line": 1882,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1882"
              },
              "summary": "Softmax using the float-native API signature."
            },
            {
              "description": "Softmax using the float-native API signature.",
              "examples": [],
              "id": "arm_softmax_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_softmax_f16",
              "params": [
                {
                  "description": "Pointer to the input matrix stored as `num_rows` rows of `row_size` values.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float16_t *"
                },
                {
                  "description": "Number of rows in the input matrix.",
                  "direction": "in",
                  "name": "num_rows",
                  "type": "int32_t"
                },
                {
                  "description": "Number of columns per row.",
                  "direction": "in",
                  "name": "row_size",
                  "type": "int32_t"
                },
                {
                  "description": "Pointer to the output matrix.",
                  "direction": "out",
                  "name": "output",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_softmax_f16(const float16_t *input, int32_t num_rows, int32_t row_size, float16_t *output)",
              "source": {
                "line": 3640,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3640"
              },
              "summary": "Softmax using the float-native API signature."
            }
          ]
        },
        {
          "description": "",
          "name": "Slicing Functions:",
          "path": "heliaCORE.StridedSlice",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Strided slice for float32 data (pure copy, TensorFlow Lite compatible).",
              "examples": [],
              "id": "arm_strided_slice_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_strided_slice_f32",
              "params": [
                {
                  "description": "Pointer to input tensor.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Pointer to output tensor.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Begin dimensions for slicing.",
                  "direction": "in",
                  "name": "begin_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Stride dimensions for slicing.",
                  "direction": "in",
                  "name": "stride_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *const"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_strided_slice_f32(\n    const float32_t *input_data,\n    float32_t *output_data,\n    const cmsis_nn_dims *const input_dims,\n    const cmsis_nn_dims *const begin_dims,\n    const cmsis_nn_dims *const stride_dims,\n    const cmsis_nn_dims *const output_dims\n)",
              "source": {
                "line": 823,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L823"
              },
              "summary": "Strided slice for float32 data (pure copy, TensorFlow Lite compatible)."
            }
          ]
        },
        {
          "description": "// end group groupPrivTypes\n\nPerform data type conversion in-between neural network operations",
          "name": "Data Conversion",
          "path": "heliaCORE.supportConversion",
          "submodules": [],
          "summary": "// end group groupPrivTypes",
          "symbols": []
        },
        {
          "description": "",
          "name": "SVDF Functions",
          "path": "heliaCORE.SVDF",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Stateful singular value decomposition filter.",
              "examples": [],
              "id": "arm_svdf_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_f32",
              "params": [
                {
                  "description": "Unused by this function. Reserved for future use; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Mandatory, not optional: a NULL input_ctx, or a NULL input_ctx->buf, is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Staging buffer written by this function, holding one element per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Sized by arm_svdf_f32_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(float32_t) bytes. Setting input_ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "input_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Mandatory, not optional: a NULL output_ctx, or a NULL output_ctx->buf, is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Staging buffer written by this function, holding one element per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Sized by arm_svdf_f32_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(float32_t) bytes, truncating division. Setting output_ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "SVDF operator parameters.",
                  "direction": "in",
                  "name": "svdf_params",
                  "type": "const cmsis_nn_svdf_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "State tensor dimensions.",
                  "direction": "in",
                  "name": "state_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the mutable state tensor.",
                  "direction": "inout",
                  "name": "state_data",
                  "type": "float32_t *"
                },
                {
                  "description": "Feature-weight tensor dimensions.",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the feature-weight tensor.",
                  "direction": "in",
                  "name": "weights_feature_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Time-weight tensor dimensions.",
                  "direction": "in",
                  "name": "weights_time_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the time-weight tensor.",
                  "direction": "in",
                  "name": "weights_time_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float32_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_svdf_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *input_ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_svdf_params_f32 *svdf_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *state_dims,\n    float32_t *state_data,\n    const cmsis_nn_dims *weights_feature_dims,\n    const float32_t *weights_feature_data,\n    const cmsis_nn_dims *weights_time_dims,\n    const float32_t *weights_time_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
              "source": {
                "line": 1694,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1694"
              },
              "summary": "Stateful singular value decomposition filter."
            },
            {
              "description": "Get size of the input_ctx staging buffer required by `arm_svdf_f32()`.\n\n:::note\nThis query reports an out-of-range shape as -1, following the SVDF family (`arm_svdf_s8_get_buffer_size()`). That differs from the float convolution and fully-connected queries in this header, which report an out-of-range size as 0. The reason is that `arm_svdf_f32()` reads ctx->size, and size == 0 is the opt-out signal for its scratch-size check: a 0-on-overflow answer fed straight back as `buf = alloc(0), size = 0` would silently disable the check over a zero-byte allocation, whereas alloc((size_t)-1) fails and the NULL check catches it.\n\n:::\n\n:::note\n0 is still a valid *return* for a degenerate shape (input_dims->n == 0). Unlike the general rule in README.md, a 0 here does NOT mean you may pass { NULL, 0 }: `arm_svdf_f32()` rejects a NULL input_ctx->buf with ARM_CMSIS_NN_ARG_ERROR regardless of the size. Allocate a non-NULL pointer, or do not call the kernel for a shape that produces no output.\n\n:::",
              "examples": [],
              "id": "arm_svdf_f32_input_ctx_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_f32_input_ctx_get_buffer_size",
              "params": [
                {
                  "description": "Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f32()`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Feature-weight tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f32()`.",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: input_dims->n * weights_feature_dims->n * sizeof(float32_t). Returns -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the product would not fit in an int32_t. The figure and the validation are the same on every build target, since `arm_svdf_f32()` stages this buffer on every build rather than only under MVE."
                }
              ],
              "signature": "int32_t arm_svdf_f32_input_ctx_get_buffer_size(\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *weights_feature_dims\n)",
              "source": {
                "line": 1734,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1734"
              },
              "summary": "Get size of the inputctx staging buffer required by armsvdff32()."
            },
            {
              "description": "Get size of the output_ctx staging buffer required by `arm_svdf_f32()`.\n\n:::note\nSame -1 and degenerate-0 contract as `arm_svdf_f32_input_ctx_get_buffer_size()`, including that a 0 does not license passing { NULL, 0 }.\n\n:::",
              "examples": [],
              "id": "arm_svdf_f32_output_ctx_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_f32_output_ctx_get_buffer_size",
              "params": [
                {
                  "description": "SVDF operator parameters; only svdf_params->rank is read.",
                  "direction": "in",
                  "name": "svdf_params",
                  "type": "const cmsis_nn_svdf_params_f32 *"
                },
                {
                  "description": "Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f32()`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Feature-weight tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f32()`.",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(float32_t), truncating division to match the kernel's own unit count. Returns -1 if any pointer is NULL, if svdf_params->rank is zero or negative, if input_dims->n or weights_feature_dims->n is negative, or if the product would not fit in an int32_t."
                }
              ],
              "signature": "int32_t arm_svdf_f32_output_ctx_get_buffer_size(\n    const cmsis_nn_svdf_params_f32 *svdf_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *weights_feature_dims\n)",
              "source": {
                "line": 1754,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1754"
              },
              "summary": "Get size of the outputctx staging buffer required by armsvdff32()."
            },
            {
              "description": "Stateful singular value decomposition filter, float16 variant.\n\n:::note\nSizing an f16 layer with the `arm_svdf_f32()` queries over-allocates and is safe. Sizing an f32 layer with the `arm_svdf_f16()` queries under-allocates by half: `arm_svdf_f32()` returns ARM_CMSIS_NN_ARG_ERROR if ctx->size carries that undersized figure, but corrupts memory if ctx->size is left at 0, which opts out of the check.\n\n:::\n\n:::note\nNaN propagates through the activation clamps that take the bit-classified scalar clamp of #380, at every optimization level on the gated toolchains, including the shipped -Ofast: the input-activation clamp is that scalar clamp on EVERY build, and the output-activation clamp is on non-MVE builds. On MVE builds the output-activation clamp is vmaxnmq/vminnmq with no NaN restore, so a NaN resolves to a clamp bound there instead.\n\n:::",
              "examples": [],
              "id": "arm_svdf_f16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_f16",
              "params": [
                {
                  "description": "Unused by this function. Reserved for future use; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Mandatory, not optional: a NULL input_ctx, or a NULL input_ctx->buf, is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Staging buffer written by this function, holding one element per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Sized by arm_svdf_f16_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(float16_t) bytes. Note this is float16_t, half the `arm_svdf_f32()` figure for the same shape. Setting input_ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "input_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Mandatory, not optional: a NULL output_ctx, or a NULL output_ctx->buf, is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Staging buffer written by this function, holding one element per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Sized by arm_svdf_f16_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(float16_t) bytes, truncating division. Note this is float16_t, half the `arm_svdf_f32()` figure for the same shape. Setting output_ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "SVDF operator parameters.",
                  "direction": "in",
                  "name": "svdf_params",
                  "type": "const cmsis_nn_svdf_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "State tensor dimensions.",
                  "direction": "in",
                  "name": "state_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the mutable state tensor.",
                  "direction": "inout",
                  "name": "state_data",
                  "type": "float16_t *"
                },
                {
                  "description": "Feature-weight tensor dimensions.",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the feature-weight tensor.",
                  "direction": "in",
                  "name": "weights_feature_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Time-weight tensor dimensions.",
                  "direction": "in",
                  "name": "weights_time_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the time-weight tensor.",
                  "direction": "in",
                  "name": "weights_time_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Bias tensor dimensions.",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias tensor data.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const float16_t *"
                },
                {
                  "description": "Output tensor dimensions.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor data.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "float16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_svdf_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *input_ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_svdf_params_f16 *svdf_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *state_dims,\n    float16_t *state_data,\n    const cmsis_nn_dims *weights_feature_dims,\n    const float16_t *weights_feature_data,\n    const cmsis_nn_dims *weights_time_dims,\n    const float16_t *weights_time_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
              "source": {
                "line": 3479,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3479"
              },
              "summary": "Stateful singular value decomposition filter, float16 variant."
            },
            {
              "description": "Get size of the input_ctx staging buffer required by `arm_svdf_f16()`.\n\n:::note\n`arm_svdf_f16()` stages float16_t, so this is HALF the byte count `arm_svdf_f32_input_ctx_get_buffer_size()` returns for the same shape. Sizing an f16 layer with the f32 query over-allocates and is safe; sizing an f32 layer with this one under-allocates by half.\n\n:::\n\n:::note\nThis query reports an out-of-range shape as -1, following the SVDF family (`arm_svdf_s8_get_buffer_size()`), not the 0 used by the float convolution and fully-connected queries in this header. The reason is that `arm_svdf_f16()` reads ctx->size, and size == 0 is the opt-out signal for its scratch-size check: a 0-on-overflow answer fed straight back as `buf = alloc(0), size = 0` would silently disable the check over a zero-byte allocation, whereas alloc((size_t)-1) fails and the NULL check catches it.\n\n:::\n\n:::note\n0 is still a valid *return* for a degenerate shape (input_dims->n == 0). Unlike the general rule in README.md, a 0 here does NOT mean you may pass { NULL, 0 }: `arm_svdf_f16()` rejects a NULL input_ctx->buf with ARM_CMSIS_NN_ARG_ERROR regardless of the size.\n\n:::",
              "examples": [],
              "id": "arm_svdf_f16_input_ctx_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_f16_input_ctx_get_buffer_size",
              "params": [
                {
                  "description": "Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f16()`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Feature-weight tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f16()`.",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: input_dims->n * weights_feature_dims->n * sizeof(float16_t). Returns -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the product would not fit in an int32_t. The figure and the validation are the same on every build target, since `arm_svdf_f16()` stages this buffer on every build rather than only under MVE."
                }
              ],
              "signature": "int32_t arm_svdf_f16_input_ctx_get_buffer_size(\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *weights_feature_dims\n)",
              "source": {
                "line": 3521,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3521"
              },
              "summary": "Get size of the inputctx staging buffer required by armsvdff16()."
            },
            {
              "description": "Get size of the output_ctx staging buffer required by `arm_svdf_f16()`.\n\n:::note\n`arm_svdf_f16()` stages float16_t, so this is HALF the byte count `arm_svdf_f32_output_ctx_get_buffer_size()` returns for the same shape.\n\n:::\n\n:::note\nSame -1 and degenerate-0 contract as `arm_svdf_f16_input_ctx_get_buffer_size()`, including that a 0 does not license passing { NULL, 0 }. A rank greater than weights_feature_dims->n truncates the unit count to 0 and so returns 0.\n\n:::",
              "examples": [],
              "id": "arm_svdf_f16_output_ctx_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_f16_output_ctx_get_buffer_size",
              "params": [
                {
                  "description": "SVDF operator parameters; only svdf_params->rank is read.",
                  "direction": "in",
                  "name": "svdf_params",
                  "type": "const cmsis_nn_svdf_params_f16 *"
                },
                {
                  "description": "Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f16()`.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Feature-weight tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_f16()`.",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(float16_t), truncating division to match the kernel's own unit count. Returns -1 if any pointer is NULL, if svdf_params->rank is zero or negative, if input_dims->n or weights_feature_dims->n is negative, or if the product would not fit in an int32_t."
                }
              ],
              "signature": "int32_t arm_svdf_f16_output_ctx_get_buffer_size(\n    const cmsis_nn_svdf_params_f16 *svdf_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *weights_feature_dims\n)",
              "source": {
                "line": 3544,
                "path": "Include/arm_nnfunctions_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3544"
              },
              "summary": "Get size of the outputctx staging buffer required by armsvdff16()."
            }
          ]
        },
        {
          "description": "",
          "name": "Internal",
          "path": "heliaCORE.Internal",
          "submodules": [
            {
              "description": "",
              "name": "arm_concatenation_common.h",
              "path": "heliaCORE.Internal.arm_concatenation_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_CONCATENATION_DEFINE",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CONCATENATION_DEFINE",
                  "params": [
                    {
                      "description": "",
                      "name": "SUFFIX"
                    },
                    {
                      "description": "",
                      "name": "TYPE"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_CONCATENATION_DEFINE(SUFFIX, TYPE)",
                  "source": {
                    "line": 36,
                    "path": "Include/Internal/arm_concatenation_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_concatenation_common.h#L36"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_conv_opt_common.h",
              "path": "heliaCORE.Internal.arm_conv_opt_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_CONV_SPEC_ENTRY",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CONV_SPEC_ENTRY",
                  "params": [
                    {
                      "description": "",
                      "name": "MATCH_FN"
                    },
                    {
                      "description": "",
                      "name": "CALL_FN"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_CONV_SPEC_ENTRY(MATCH_FN, CALL_FN) {(MATCH_FN), (CALL_FN)}",
                  "source": {
                    "line": 34,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L34"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_CONV_ARRAY_SIZE",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CONV_ARRAY_SIZE",
                  "params": [
                    {
                      "description": "",
                      "name": "arr"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_CONV_ARRAY_SIZE(arr) (sizeof(arr) / sizeof((arr)[0]))",
                  "source": {
                    "line": 36,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L36"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MAX_TILE_ROWS",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MAX_TILE_ROWS",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MAX_TILE_ROWS (8)",
                  "source": {
                    "line": 45,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L45"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_K",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_K",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_K (16)",
                  "source": {
                    "line": 46,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L46"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_OC",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_OC",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_OC (8)",
                  "source": {
                    "line": 47,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L47"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_POS",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_POS",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_CONV_NHWC_PATCH_GEMM_F32_MIN_POS (8)",
                  "source": {
                    "line": 48,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L48"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MAX_TILE_ROWS",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MAX_TILE_ROWS",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MAX_TILE_ROWS (8)",
                  "source": {
                    "line": 54,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L54"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_K",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_K",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_K (16)",
                  "source": {
                    "line": 55,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L55"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_OC",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_OC",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_OC (8)",
                  "source": {
                    "line": 56,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L56"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_POS",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_POS",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_CONV_NHWC_PATCH_GEMM_F16_MIN_POS (8)",
                  "source": {
                    "line": 57,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L57"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_CONV_DISPATCH",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CONV_DISPATCH",
                  "params": [
                    {
                      "description": "",
                      "name": "TABLE"
                    },
                    {
                      "description": "",
                      "name": "COUNT"
                    },
                    {
                      "description": "",
                      "name": "..."
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_CONV_DISPATCH(TABLE, COUNT, ...) do \\ { \\ for (size_t _i = 0; _i < (COUNT); ++_i) \\ { \\ if ((TABLE)[_i].match(__VA_ARGS__)) \\ { \\ return (TABLE)[_i].call(__VA_ARGS__); \\ } \\ } \\ } while (0)",
                  "source": {
                    "line": 59,
                    "path": "Include/Internal/arm_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_common.h#L59"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_conv_opt_f16.h",
              "path": "heliaCORE.Internal.arm_conv_opt_f16",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_spec_f16",
                  "kind": "struct",
                  "language": "c",
                  "members": [
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_conv_spec_f16::match",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "match",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_conv_match_f16 match",
                      "source": {
                        "line": 65,
                        "path": "Include/Internal/arm_conv_opt_f16.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L65"
                      },
                      "summary": ""
                    },
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_conv_spec_f16::call",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "call",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_conv_call_f16 call",
                      "source": {
                        "line": 66,
                        "path": "Include/Internal/arm_conv_opt_f16.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L66"
                      },
                      "summary": ""
                    }
                  ],
                  "name": "arm_conv_spec_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "struct arm_conv_spec_f16",
                  "source": {
                    "line": 63,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L63"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_match_f16",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv_match_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef bool(* arm_conv_match_f16) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input_data, const cmsis_nn_dims *filter_dims, const float16_t *filter_data, const cmsis_nn_dims *bias_dims, const float16_t *bias_data, const cmsis_nn_dims *output_dims, float16_t *output_data)",
                  "source": {
                    "line": 41,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L41"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_call_f16",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv_call_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef arm_cmsis_nn_status(* arm_conv_call_f16) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input_data, const cmsis_nn_dims *filter_dims, const float16_t *filter_data, const cmsis_nn_dims *bias_dims, const float16_t *bias_data, const cmsis_nn_dims *output_dims, float16_t *output_data)",
                  "source": {
                    "line": 52,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L52"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k5_nhwc_f16_match",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k5_nhwc_f16_match",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float16_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_conv1d_spec_k5_nhwc_f16_match(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
                  "source": {
                    "line": 69,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L69"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k5_nhwc_f16_call_body",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k5_nhwc_f16_call_body",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float16_t *"
                    },
                    {
                      "description": "",
                      "name": "acc16",
                      "type": "const bool"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_conv1d_spec_k5_nhwc_f16_call_body(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    const bool acc16\n)",
                  "source": {
                    "line": 103,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L103"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k5_nhwc_f16_call",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k5_nhwc_f16_call",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float16_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_conv1d_spec_k5_nhwc_f16_call(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
                  "source": {
                    "line": 141,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L141"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k5_nhwc_f16_call_acc16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k5_nhwc_f16_call_acc16",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float16_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_conv1d_spec_k5_nhwc_f16_call_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
                  "source": {
                    "line": 165,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L165"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k3_nhwc_f16_match",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k3_nhwc_f16_match",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float16_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_conv1d_spec_k3_nhwc_f16_match(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
                  "source": {
                    "line": 189,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L189"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k3_nhwc_f16_call_body",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k3_nhwc_f16_call_body",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float16_t *"
                    },
                    {
                      "description": "",
                      "name": "acc16",
                      "type": "const bool"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_conv1d_spec_k3_nhwc_f16_call_body(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    const bool acc16\n)",
                  "source": {
                    "line": 223,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L223"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k3_nhwc_f16_call",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k3_nhwc_f16_call",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float16_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_conv1d_spec_k3_nhwc_f16_call(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
                  "source": {
                    "line": 261,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L261"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k3_nhwc_f16_call_acc16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k3_nhwc_f16_call_acc16",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float16_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_conv1d_spec_k3_nhwc_f16_call_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
                  "source": {
                    "line": 285,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L285"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_spec_nhwc_f16",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv_spec_nhwc_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const arm_conv_spec_f16 arm_conv_spec_nhwc_f16[]",
                  "source": {
                    "line": 310,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L310"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_spec_nhwc_f16_acc16",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv_spec_nhwc_f16_acc16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const arm_conv_spec_f16 arm_conv_spec_nhwc_f16_acc16[]",
                  "source": {
                    "line": 317,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L317"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_spec_nhwc_f16_matches_any",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv_spec_nhwc_f16_matches_any",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_conv_spec_nhwc_f16_matches_any(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
                  "source": {
                    "line": 326,
                    "path": "Include/Internal/arm_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f16.h#L326"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_conv_opt_f32.h",
              "path": "heliaCORE.Internal.arm_conv_opt_f32",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_spec_f32",
                  "kind": "struct",
                  "language": "c",
                  "members": [
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_conv_spec_f32::match",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "match",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_conv_match_f32 match",
                      "source": {
                        "line": 65,
                        "path": "Include/Internal/arm_conv_opt_f32.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L65"
                      },
                      "summary": ""
                    },
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_conv_spec_f32::call",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "call",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_conv_call_f32 call",
                      "source": {
                        "line": 66,
                        "path": "Include/Internal/arm_conv_opt_f32.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L66"
                      },
                      "summary": ""
                    }
                  ],
                  "name": "arm_conv_spec_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "struct arm_conv_spec_f32",
                  "source": {
                    "line": 63,
                    "path": "Include/Internal/arm_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L63"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_match_f32",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv_match_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef bool(* arm_conv_match_f32) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input_data, const cmsis_nn_dims *filter_dims, const float32_t *filter_data, const cmsis_nn_dims *bias_dims, const float32_t *bias_data, const cmsis_nn_dims *output_dims, float32_t *output_data)",
                  "source": {
                    "line": 41,
                    "path": "Include/Internal/arm_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L41"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_call_f32",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv_call_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef arm_cmsis_nn_status(* arm_conv_call_f32) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input_data, const cmsis_nn_dims *filter_dims, const float32_t *filter_data, const cmsis_nn_dims *bias_dims, const float32_t *bias_data, const cmsis_nn_dims *output_dims, float32_t *output_data)",
                  "source": {
                    "line": 52,
                    "path": "Include/Internal/arm_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L52"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k5_nhwc_f32_match",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k5_nhwc_f32_match",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f32 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float32_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_conv1d_spec_k5_nhwc_f32_match(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
                  "source": {
                    "line": 69,
                    "path": "Include/Internal/arm_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L69"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k5_nhwc_f32_call",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k5_nhwc_f32_call",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f32 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float32_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_conv1d_spec_k5_nhwc_f32_call(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
                  "source": {
                    "line": 103,
                    "path": "Include/Internal/arm_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L103"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k3_nhwc_f32_match",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k3_nhwc_f32_match",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f32 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float32_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_conv1d_spec_k3_nhwc_f32_match(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
                  "source": {
                    "line": 140,
                    "path": "Include/Internal/arm_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L140"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1d_spec_k3_nhwc_f32_call",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1d_spec_k3_nhwc_f32_call",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f32 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "float32_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_conv1d_spec_k3_nhwc_f32_call(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
                  "source": {
                    "line": 174,
                    "path": "Include/Internal/arm_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L174"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_spec_nhwc_f32",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv_spec_nhwc_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const arm_conv_spec_f32 arm_conv_spec_nhwc_f32[]",
                  "source": {
                    "line": 211,
                    "path": "Include/Internal/arm_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L211"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv_spec_nhwc_f32_matches_any",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv_spec_nhwc_f32_matches_any",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_conv_params_f32 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_conv_spec_nhwc_f32_matches_any(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
                  "source": {
                    "line": 216,
                    "path": "Include/Internal/arm_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv_opt_f32.h#L216"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_conv1x1_opt_common.h",
              "path": "heliaCORE.Internal.arm_conv1x1_opt_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_CONV1X1_SPEC_ENTRY",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CONV1X1_SPEC_ENTRY",
                  "params": [
                    {
                      "description": "",
                      "name": "MATCH_FN"
                    },
                    {
                      "description": "",
                      "name": "CALL_FN"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_CONV1X1_SPEC_ENTRY(MATCH_FN, CALL_FN) { \\ (MATCH_FN), (CALL_FN) \\ }",
                  "source": {
                    "line": 34,
                    "path": "Include/Internal/arm_conv1x1_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_common.h#L34"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_CONV1X1_ARRAY_SIZE",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CONV1X1_ARRAY_SIZE",
                  "params": [
                    {
                      "description": "",
                      "name": "arr"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_CONV1X1_ARRAY_SIZE(arr) (sizeof(arr) / sizeof((arr)[0]))",
                  "source": {
                    "line": 39,
                    "path": "Include/Internal/arm_conv1x1_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_common.h#L39"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_CONV1X1_DISPATCH",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_CONV1X1_DISPATCH",
                  "params": [
                    {
                      "description": "",
                      "name": "TABLE"
                    },
                    {
                      "description": "",
                      "name": "COUNT"
                    },
                    {
                      "description": "",
                      "name": "..."
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_CONV1X1_DISPATCH(TABLE, COUNT, ...) do \\ { \\ for (size_t _i = 0; _i < (COUNT); ++_i) \\ { \\ if ((TABLE)[_i].match(__VA_ARGS__)) \\ { \\ return (TABLE)[_i].call(__VA_ARGS__); \\ } \\ } \\ } while (0)",
                  "source": {
                    "line": 41,
                    "path": "Include/Internal/arm_conv1x1_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_common.h#L41"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_conv1x1_opt_f16.h",
              "path": "heliaCORE.Internal.arm_conv1x1_opt_f16",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1x1_spec_f16",
                  "kind": "struct",
                  "language": "c",
                  "members": [
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_conv1x1_spec_f16::match",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "match",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_conv1x1_match_f16 match",
                      "source": {
                        "line": 64,
                        "path": "Include/Internal/arm_conv1x1_opt_f16.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f16.h#L64"
                      },
                      "summary": ""
                    },
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_conv1x1_spec_f16::call",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "call",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_conv1x1_call_f16 call",
                      "source": {
                        "line": 65,
                        "path": "Include/Internal/arm_conv1x1_opt_f16.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f16.h#L65"
                      },
                      "summary": ""
                    }
                  ],
                  "name": "arm_conv1x1_spec_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "struct arm_conv1x1_spec_f16",
                  "source": {
                    "line": 62,
                    "path": "Include/Internal/arm_conv1x1_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f16.h#L62"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1x1_match_f16",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1x1_match_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef bool(* arm_conv1x1_match_f16) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input_data, const cmsis_nn_dims *filter_dims, const float16_t *filter_data, const cmsis_nn_dims *bias_dims, const float16_t *bias_data, const cmsis_nn_dims *output_dims, float16_t *output_data)",
                  "source": {
                    "line": 40,
                    "path": "Include/Internal/arm_conv1x1_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f16.h#L40"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1x1_call_f16",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1x1_call_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef arm_cmsis_nn_status(* arm_conv1x1_call_f16) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input_data, const cmsis_nn_dims *filter_dims, const float16_t *filter_data, const cmsis_nn_dims *bias_dims, const float16_t *bias_data, const cmsis_nn_dims *output_dims, float16_t *output_data)",
                  "source": {
                    "line": 51,
                    "path": "Include/Internal/arm_conv1x1_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f16.h#L51"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_conv1x1_opt_f32.h",
              "path": "heliaCORE.Internal.arm_conv1x1_opt_f32",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1x1_spec_f32",
                  "kind": "struct",
                  "language": "c",
                  "members": [
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_conv1x1_spec_f32::match",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "match",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_conv1x1_match_f32 match",
                      "source": {
                        "line": 64,
                        "path": "Include/Internal/arm_conv1x1_opt_f32.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f32.h#L64"
                      },
                      "summary": ""
                    },
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_conv1x1_spec_f32::call",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "call",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_conv1x1_call_f32 call",
                      "source": {
                        "line": 65,
                        "path": "Include/Internal/arm_conv1x1_opt_f32.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f32.h#L65"
                      },
                      "summary": ""
                    }
                  ],
                  "name": "arm_conv1x1_spec_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "struct arm_conv1x1_spec_f32",
                  "source": {
                    "line": 62,
                    "path": "Include/Internal/arm_conv1x1_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f32.h#L62"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1x1_match_f32",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1x1_match_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef bool(* arm_conv1x1_match_f32) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input_data, const cmsis_nn_dims *filter_dims, const float32_t *filter_data, const cmsis_nn_dims *bias_dims, const float32_t *bias_data, const cmsis_nn_dims *output_dims, float32_t *output_data)",
                  "source": {
                    "line": 40,
                    "path": "Include/Internal/arm_conv1x1_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f32.h#L40"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_conv1x1_call_f32",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_conv1x1_call_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef arm_cmsis_nn_status(* arm_conv1x1_call_f32) (const cmsis_nn_context *ctx, const cmsis_nn_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input_data, const cmsis_nn_dims *filter_dims, const float32_t *filter_data, const cmsis_nn_dims *bias_dims, const float32_t *bias_data, const cmsis_nn_dims *output_dims, float32_t *output_data)",
                  "source": {
                    "line": 51,
                    "path": "Include/Internal/arm_conv1x1_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_conv1x1_opt_f32.h#L51"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_depthwise_conv_opt_common.h",
              "path": "heliaCORE.Internal.arm_depthwise_conv_opt_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_DW_SPEC_ENTRY",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_DW_SPEC_ENTRY",
                  "params": [
                    {
                      "description": "",
                      "name": "MATCH_FN"
                    },
                    {
                      "description": "",
                      "name": "CALL_FN"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_DW_SPEC_ENTRY(MATCH_FN, CALL_FN) { \\ (MATCH_FN), (CALL_FN) \\ }",
                  "source": {
                    "line": 36,
                    "path": "Include/Internal/arm_depthwise_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_common.h#L36"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_DW_ARRAY_SIZE",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_DW_ARRAY_SIZE",
                  "params": [
                    {
                      "description": "",
                      "name": "arr"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_DW_ARRAY_SIZE(arr) (sizeof(arr) / sizeof((arr)[0]))",
                  "source": {
                    "line": 41,
                    "path": "Include/Internal/arm_depthwise_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_common.h#L41"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_depthwise_conv_input_index_nhwc",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_depthwise_conv_input_index_nhwc",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "y",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "c",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "input_x",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "input_ch",
                      "type": "int32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static int32_t arm_depthwise_conv_input_index_nhwc(int32_t x, int32_t y, int32_t c, int32_t input_x, int32_t input_ch)",
                  "source": {
                    "line": 45,
                    "path": "Include/Internal/arm_depthwise_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_common.h#L45"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_depthwise_conv_output_index_nhwc",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_depthwise_conv_output_index_nhwc",
                  "params": [
                    {
                      "description": "",
                      "name": "out_x",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "out_y",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "out_ch",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "output_x",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "output_ch",
                      "type": "int32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static int32_t arm_depthwise_conv_output_index_nhwc(\n    int32_t out_x,\n    int32_t out_y,\n    int32_t out_ch,\n    int32_t output_x,\n    int32_t output_ch\n)",
                  "source": {
                    "line": 52,
                    "path": "Include/Internal/arm_depthwise_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_common.h#L52"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_DW_DISPATCH",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_DW_DISPATCH",
                  "params": [
                    {
                      "description": "",
                      "name": "TABLE"
                    },
                    {
                      "description": "",
                      "name": "COUNT"
                    },
                    {
                      "description": "",
                      "name": "..."
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_DW_DISPATCH(TABLE, COUNT, ...) do \\ { \\ for (size_t _i = 0; _i < (COUNT); ++_i) \\ { \\ if ((TABLE)[_i].match(__VA_ARGS__)) \\ { \\ return (TABLE)[_i].call(__VA_ARGS__); \\ } \\ } \\ } while (0)",
                  "source": {
                    "line": 57,
                    "path": "Include/Internal/arm_depthwise_conv_opt_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_common.h#L57"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_depthwise_conv_opt_f16.h",
              "path": "heliaCORE.Internal.arm_depthwise_conv_opt_f16",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_f16",
                  "kind": "struct",
                  "language": "c",
                  "members": [
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_dw_spec_f16::match",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "match",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_dw_match_f16 match",
                      "source": {
                        "line": 69,
                        "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L69"
                      },
                      "summary": ""
                    },
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_dw_spec_f16::call",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "call",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_dw_call_f16 call",
                      "source": {
                        "line": 70,
                        "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L70"
                      },
                      "summary": ""
                    }
                  ],
                  "name": "arm_dw_spec_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "struct arm_dw_spec_f16",
                  "source": {
                    "line": 67,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L67"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_match_f16",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_match_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef bool(* arm_dw_match_f16) (const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input, const cmsis_nn_dims *filter_dims, const float16_t *kernel, const cmsis_nn_dims *bias_dims, const float16_t *bias, const cmsis_nn_dims *output_dims, float16_t *output, arm_nn_dw_kernel_layout_f16 kernel_layout)",
                  "source": {
                    "line": 43,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L43"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_call_f16",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_call_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef arm_cmsis_nn_status(* arm_dw_call_f16) (const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params_f16 *params, const cmsis_nn_dims *input_dims, const float16_t *input, const cmsis_nn_dims *filter_dims, const float16_t *kernel, const cmsis_nn_dims *bias_dims, const float16_t *bias, const cmsis_nn_dims *output_dims, float16_t *output, arm_nn_dw_kernel_layout_f16 kernel_layout)",
                  "source": {
                    "line": 55,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L55"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_k3_1d_nhwc_f16_match",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_spec_k3_1d_nhwc_f16_match",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_dw_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "kernel",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output",
                      "type": "float16_t *"
                    },
                    {
                      "description": "",
                      "name": "kernel_layout",
                      "type": "arm_nn_dw_kernel_layout_f16"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_dw_spec_k3_1d_nhwc_f16_match(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_dw_kernel_layout_f16 kernel_layout\n)",
                  "source": {
                    "line": 73,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L73"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_k3_1d_nhwc_f16_call",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_spec_k3_1d_nhwc_f16_call",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_dw_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "kernel",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output",
                      "type": "float16_t *"
                    },
                    {
                      "description": "",
                      "name": "kernel_layout",
                      "type": "arm_nn_dw_kernel_layout_f16"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_dw_spec_k3_1d_nhwc_f16_call(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_dw_kernel_layout_f16 kernel_layout\n)",
                  "source": {
                    "line": 105,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L105"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_2x5_nhwc_f16_match",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_spec_2x5_nhwc_f16_match",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_dw_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "kernel",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output",
                      "type": "float16_t *"
                    },
                    {
                      "description": "",
                      "name": "kernel_layout",
                      "type": "arm_nn_dw_kernel_layout_f16"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_dw_spec_2x5_nhwc_f16_match(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_dw_kernel_layout_f16 kernel_layout\n)",
                  "source": {
                    "line": 134,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L134"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_2x5_nhwc_f16_call",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_spec_2x5_nhwc_f16_call",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_dw_conv_params_f16 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "kernel",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output",
                      "type": "float16_t *"
                    },
                    {
                      "description": "",
                      "name": "kernel_layout",
                      "type": "arm_nn_dw_kernel_layout_f16"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_dw_spec_2x5_nhwc_f16_call(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_dw_kernel_layout_f16 kernel_layout\n)",
                  "source": {
                    "line": 163,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L163"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_nhwc_f16",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_spec_nhwc_f16",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const arm_dw_spec_f16 arm_dw_spec_nhwc_f16[]",
                  "source": {
                    "line": 199,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f16.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f16.h#L199"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_depthwise_conv_opt_f32.h",
              "path": "heliaCORE.Internal.arm_depthwise_conv_opt_f32",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_f32",
                  "kind": "struct",
                  "language": "c",
                  "members": [
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_dw_spec_f32::match",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "match",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_dw_match_f32 match",
                      "source": {
                        "line": 69,
                        "path": "Include/Internal/arm_depthwise_conv_opt_f32.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f32.h#L69"
                      },
                      "summary": ""
                    },
                    {
                      "description": "",
                      "examples": [],
                      "id": "arm_dw_spec_f32::call",
                      "kind": "attribute",
                      "language": "c",
                      "members": [],
                      "name": "call",
                      "params": [],
                      "raises": [],
                      "returns": [],
                      "signature": "arm_dw_call_f32 call",
                      "source": {
                        "line": 70,
                        "path": "Include/Internal/arm_depthwise_conv_opt_f32.h",
                        "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f32.h#L70"
                      },
                      "summary": ""
                    }
                  ],
                  "name": "arm_dw_spec_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "struct arm_dw_spec_f32",
                  "source": {
                    "line": 67,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f32.h#L67"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_match_f32",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_match_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef bool(* arm_dw_match_f32) (const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input, const cmsis_nn_dims *filter_dims, const float32_t *kernel, const cmsis_nn_dims *bias_dims, const float32_t *bias, const cmsis_nn_dims *output_dims, float32_t *output, arm_nn_dw_kernel_layout_f32 kernel_layout)",
                  "source": {
                    "line": 43,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f32.h#L43"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_call_f32",
                  "kind": "type",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_call_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "typedef arm_cmsis_nn_status(* arm_dw_call_f32) (const cmsis_nn_context *ctx, const cmsis_nn_dw_conv_params_f32 *params, const cmsis_nn_dims *input_dims, const float32_t *input, const cmsis_nn_dims *filter_dims, const float32_t *kernel, const cmsis_nn_dims *bias_dims, const float32_t *bias, const cmsis_nn_dims *output_dims, float32_t *output, arm_nn_dw_kernel_layout_f32 kernel_layout)",
                  "source": {
                    "line": 55,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f32.h#L55"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_k3_1d_nhwc_f32_match",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_spec_k3_1d_nhwc_f32_match",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_dw_conv_params_f32 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "kernel",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output",
                      "type": "float32_t *"
                    },
                    {
                      "description": "",
                      "name": "kernel_layout",
                      "type": "arm_nn_dw_kernel_layout_f32"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_dw_spec_k3_1d_nhwc_f32_match(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f32 *params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output,\n    arm_nn_dw_kernel_layout_f32 kernel_layout\n)",
                  "source": {
                    "line": 73,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f32.h#L73"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_k3_1d_nhwc_f32_call",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_spec_k3_1d_nhwc_f32_call",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "params",
                      "type": "const cmsis_nn_dw_conv_params_f32 *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "input",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "filter_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "kernel",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "bias_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "bias",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output",
                      "type": "float32_t *"
                    },
                    {
                      "description": "",
                      "name": "kernel_layout",
                      "type": "arm_nn_dw_kernel_layout_f32"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_dw_spec_k3_1d_nhwc_f32_call(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f32 *params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output,\n    arm_nn_dw_kernel_layout_f32 kernel_layout\n)",
                  "source": {
                    "line": 105,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f32.h#L105"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_dw_spec_nhwc_f32",
                  "kind": "attribute",
                  "language": "c",
                  "members": [],
                  "name": "arm_dw_spec_nhwc_f32",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "const arm_dw_spec_f32 arm_dw_spec_nhwc_f32[]",
                  "source": {
                    "line": 135,
                    "path": "Include/Internal/arm_depthwise_conv_opt_f32.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_depthwise_conv_opt_f32.h#L135"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_minmax_f16_common.h",
              "path": "heliaCORE.Internal.arm_minmax_f16_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "Shared implementation for arm_minimum_f16 / arm_maximum_f16.\n\nDefined out-of-line in arm_minmax_f16_common.c (rather than inline in this header) so that source-level test coverage tools (e.g. gcov) attribute hits/lines/branches to the real implementation instead of collapsing them into the call site of the thin wrapper that invokes it.",
                  "examples": [],
                  "id": "arm_minmax_f16_impl",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_minmax_f16_impl",
                  "params": [
                    {
                      "description": "Function context. Unused; may be NULL.",
                      "direction": "in",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "Pointer to the first input tensor data.",
                      "direction": "in",
                      "name": "input_1_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "Dimensions of the first input tensor.",
                      "direction": "in",
                      "name": "input_1_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "Pointer to the second input tensor data.",
                      "direction": "in",
                      "name": "input_2_data",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "Dimensions of the second input tensor.",
                      "direction": "in",
                      "name": "input_2_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "Pointer to the output tensor data.",
                      "direction": "out",
                      "name": "output_data",
                      "type": "float16_t *"
                    },
                    {
                      "description": "Dimensions of the output tensor.",
                      "direction": "in",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "Non-zero to compute elementwise maximum, zero for minimum.",
                      "direction": "in",
                      "name": "select_max",
                      "type": "int32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [
                    {
                      "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                    }
                  ],
                  "signature": "arm_cmsis_nn_status arm_minmax_f16_impl(\n    const cmsis_nn_context *ctx,\n    const float16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    int32_t select_max\n)",
                  "source": {
                    "line": 59,
                    "path": "Include/Internal/arm_minmax_f16_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_minmax_f16_common.h#L59"
                  },
                  "summary": "Shared implementation for armminimumf16 / armmaximumf16."
                }
              ]
            },
            {
              "description": "",
              "name": "arm_minmax_f32_common.h",
              "path": "heliaCORE.Internal.arm_minmax_f32_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "Shared implementation for arm_minimum_f32 / arm_maximum_f32.\n\nDefined out-of-line in arm_minmax_f32_common.c (rather than inline in this header) so that source-level test coverage tools (e.g. gcov) attribute hits/lines/branches to the real implementation instead of collapsing them into the call site of the thin wrapper that invokes it.",
                  "examples": [],
                  "id": "arm_minmax_f32_impl",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_minmax_f32_impl",
                  "params": [
                    {
                      "description": "Function context. Unused; may be NULL.",
                      "direction": "in",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "Pointer to the first input tensor data.",
                      "direction": "in",
                      "name": "input_1_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "Dimensions of the first input tensor.",
                      "direction": "in",
                      "name": "input_1_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "Pointer to the second input tensor data.",
                      "direction": "in",
                      "name": "input_2_data",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "Dimensions of the second input tensor.",
                      "direction": "in",
                      "name": "input_2_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "Pointer to the output tensor data.",
                      "direction": "out",
                      "name": "output_data",
                      "type": "float32_t *"
                    },
                    {
                      "description": "Dimensions of the output tensor.",
                      "direction": "in",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "Non-zero to compute elementwise maximum, zero for minimum.",
                      "direction": "in",
                      "name": "select_max",
                      "type": "int32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [
                    {
                      "description": "ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                    }
                  ],
                  "signature": "arm_cmsis_nn_status arm_minmax_f32_impl(\n    const cmsis_nn_context *ctx,\n    const float32_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const float32_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    float32_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    int32_t select_max\n)",
                  "source": {
                    "line": 59,
                    "path": "Include/Internal/arm_minmax_f32_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_minmax_f32_common.h#L59"
                  },
                  "summary": "Shared implementation for armminimumf32 / armmaximumf32."
                }
              ]
            },
            {
              "description": "",
              "name": "arm_nn_activation_flt.h",
              "path": "heliaCORE.Internal.arm_nn_activation_flt",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_TANH_F32_XMAX",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_TANH_F32_XMAX",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_TANH_F32_XMAX (6.0f)",
                  "source": {
                    "line": 48,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L48"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_TANH_F32_LUT_SEGMENTS",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_TANH_F32_LUT_SEGMENTS",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_TANH_F32_LUT_SEGMENTS (384)",
                  "source": {
                    "line": 49,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L49"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_TANH_F32_LUT_MAX_IDX",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_TANH_F32_LUT_MAX_IDX",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_TANH_F32_LUT_MAX_IDX (ARM_NN_TANH_F32_LUT_SEGMENTS - 1)",
                  "source": {
                    "line": 50,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L50"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_hardswish_scalar_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_hardswish_scalar_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float32_t arm_nn_hardswish_scalar_f32(float32_t x)",
                  "source": {
                    "line": 52,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L52"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_tanh_scalar_ref_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_tanh_scalar_ref_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float32_t arm_nn_tanh_scalar_ref_f32(float32_t x)",
                  "source": {
                    "line": 92,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L92"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_sigmoid_scalar_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_sigmoid_scalar_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float32_t arm_nn_sigmoid_scalar_f32(float32_t x)",
                  "source": {
                    "line": 177,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L177"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_propagate_nan_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_propagate_nan_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float32_t"
                    },
                    {
                      "description": "",
                      "name": "y",
                      "type": "float32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float32_t arm_nn_propagate_nan_f32(float32_t x, float32_t y)",
                  "source": {
                    "line": 202,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L202"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_apply_activation_type_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_apply_activation_type_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float32_t"
                    },
                    {
                      "description": "",
                      "name": "type",
                      "type": "arm_nn_activation_type_flt"
                    },
                    {
                      "description": "",
                      "name": "act_param",
                      "type": "float32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float32_t arm_nn_apply_activation_type_f32(float32_t x, arm_nn_activation_type_flt type, float32_t act_param)",
                  "source": {
                    "line": 216,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L216"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_clamp_scalar_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_clamp_scalar_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float32_t"
                    },
                    {
                      "description": "",
                      "name": "min_v",
                      "type": "float32_t"
                    },
                    {
                      "description": "",
                      "name": "max_v",
                      "type": "float32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float32_t arm_nn_clamp_scalar_f32(float32_t x, float32_t min_v, float32_t max_v)",
                  "source": {
                    "line": 248,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L248"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_vector_clamp_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_vector_clamp_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "data",
                      "type": "float32_t *"
                    },
                    {
                      "description": "",
                      "name": "block_size",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "activation_min",
                      "type": "float32_t"
                    },
                    {
                      "description": "",
                      "name": "activation_max",
                      "type": "float32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static void arm_nn_vector_clamp_f32(\n    float32_t *data,\n    int32_t block_size,\n    float32_t activation_min,\n    float32_t activation_max\n)",
                  "source": {
                    "line": 365,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L365"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_clamp_scalar_f16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_clamp_scalar_f16",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float16_t"
                    },
                    {
                      "description": "",
                      "name": "min_v",
                      "type": "float16_t"
                    },
                    {
                      "description": "",
                      "name": "max_v",
                      "type": "float16_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float16_t arm_nn_clamp_scalar_f16(float16_t x, float16_t min_v, float16_t max_v)",
                  "source": {
                    "line": 389,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L389"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_hardswish_scalar_f16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_hardswish_scalar_f16",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float16_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float16_t arm_nn_hardswish_scalar_f16(float16_t x)",
                  "source": {
                    "line": 400,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L400"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_tanh_scalar_ref_f16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_tanh_scalar_ref_f16",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float16_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float16_t arm_nn_tanh_scalar_ref_f16(float16_t x)",
                  "source": {
                    "line": 418,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L418"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_sigmoid_scalar_f16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_sigmoid_scalar_f16",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float16_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float16_t arm_nn_sigmoid_scalar_f16(float16_t x)",
                  "source": {
                    "line": 460,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L460"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_apply_activation_type_f16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_apply_activation_type_f16",
                  "params": [
                    {
                      "description": "",
                      "name": "x",
                      "type": "float16_t"
                    },
                    {
                      "description": "",
                      "name": "type",
                      "type": "arm_nn_activation_type_flt"
                    },
                    {
                      "description": "",
                      "name": "act_param",
                      "type": "float16_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float16_t arm_nn_apply_activation_type_f16(float16_t x, arm_nn_activation_type_flt type, float16_t act_param)",
                  "source": {
                    "line": 472,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L472"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_vector_clamp_f16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_vector_clamp_f16",
                  "params": [
                    {
                      "description": "",
                      "name": "data",
                      "type": "float16_t *"
                    },
                    {
                      "description": "",
                      "name": "block_size",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "activation_min",
                      "type": "float16_t"
                    },
                    {
                      "description": "",
                      "name": "activation_max",
                      "type": "float16_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static void arm_nn_vector_clamp_f16(\n    float16_t *data,\n    int32_t block_size,\n    float16_t activation_min,\n    float16_t activation_max\n)",
                  "source": {
                    "line": 577,
                    "path": "Include/Internal/arm_nn_activation_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_activation_flt.h#L577"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_nn_arg_extrema_flt.h",
              "path": "heliaCORE.Internal.arm_nn_arg_extrema_flt",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_arg_count",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_arg_count",
                  "params": [
                    {
                      "description": "",
                      "name": "dims",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "width",
                      "type": "size_t"
                    },
                    {
                      "description": "",
                      "name": "count",
                      "type": "size_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_nn_arg_count(const int32_t dims, size_t width, size_t *count)",
                  "source": {
                    "line": 24,
                    "path": "Include/Internal/arm_nn_arg_extrema_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_arg_extrema_flt.h#L24"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_arg_load",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_arg_load",
                  "params": [
                    {
                      "description": "",
                      "name": "input",
                      "type": "const uint8_t *"
                    },
                    {
                      "description": "",
                      "name": "width",
                      "type": "size_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static uint32_t arm_nn_arg_load(const uint8_t *input, size_t width)",
                  "source": {
                    "line": 46,
                    "path": "Include/Internal/arm_nn_arg_extrema_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_arg_extrema_flt.h#L46"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_arg_key",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_arg_key",
                  "params": [
                    {
                      "description": "",
                      "name": "bits",
                      "type": "uint32_t"
                    },
                    {
                      "description": "",
                      "name": "sign",
                      "type": "uint32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static uint32_t arm_nn_arg_key(uint32_t bits, uint32_t sign)",
                  "source": {
                    "line": 61,
                    "path": "Include/Internal/arm_nn_arg_extrema_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_arg_extrema_flt.h#L61"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_arg_extrema",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_arg_extrema",
                  "params": [
                    {
                      "description": "",
                      "name": "input_data",
                      "type": "const void *"
                    },
                    {
                      "description": "",
                      "name": "input_dims",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "axis",
                      "type": "int32_t"
                    },
                    {
                      "description": "",
                      "name": "output_data",
                      "type": "int32_t *"
                    },
                    {
                      "description": "",
                      "name": "width",
                      "type": "size_t"
                    },
                    {
                      "description": "",
                      "name": "maximum",
                      "type": "bool"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_nn_arg_extrema(\n    const void *input_data,\n    const cmsis_nn_dims *input_dims,\n    int32_t axis,\n    int32_t *output_data,\n    size_t width,\n    bool maximum\n)",
                  "source": {
                    "line": 72,
                    "path": "Include/Internal/arm_nn_arg_extrema_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_arg_extrema_flt.h#L72"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_nn_axis_copy_common.h",
              "path": "heliaCORE.Internal.arm_nn_axis_copy_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_axis_copy_plan",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_axis_copy_plan",
                  "params": [
                    {
                      "description": "",
                      "name": "shape",
                      "type": "const int32_t *"
                    },
                    {
                      "description": "",
                      "name": "dims",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "axis",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "inner_begin",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "axis_len",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "num",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "sizes",
                      "type": "const int32_t *"
                    },
                    {
                      "description": "",
                      "name": "outer",
                      "type": "int32_t *"
                    },
                    {
                      "description": "",
                      "name": "inner",
                      "type": "int32_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_nn_axis_copy_plan(\n    const int32_t *shape,\n    const int32_t dims,\n    const int32_t axis,\n    const int32_t inner_begin,\n    const int32_t axis_len,\n    const int32_t num,\n    const int32_t *sizes,\n    int32_t *outer,\n    int32_t *inner\n)",
                  "source": {
                    "line": 45,
                    "path": "Include/Internal/arm_nn_axis_copy_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_axis_copy_common.h#L45"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_axis_copy_ptrs_ok",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_axis_copy_ptrs_ok",
                  "params": [
                    {
                      "description": "",
                      "name": "ptrs",
                      "type": "const void *const *"
                    },
                    {
                      "description": "",
                      "name": "num",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "total",
                      "type": "const int32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static int32_t arm_nn_axis_copy_ptrs_ok(const void *const *ptrs, const int32_t num, const int32_t total)",
                  "source": {
                    "line": 123,
                    "path": "Include/Internal/arm_nn_axis_copy_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_axis_copy_common.h#L123"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_AXIS_COPY_DEFINE",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_AXIS_COPY_DEFINE",
                  "params": [
                    {
                      "description": "",
                      "name": "SUFFIX"
                    },
                    {
                      "description": "",
                      "name": "TYPE"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_AXIS_COPY_DEFINE(SUFFIX, TYPE)",
                  "source": {
                    "line": 145,
                    "path": "Include/Internal/arm_nn_axis_copy_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_axis_copy_common.h#L145"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_axis_scatter_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_axis_scatter_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "packed",
                      "type": "const float32_t *"
                    },
                    {
                      "description": "",
                      "name": "outer",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "num",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "sizes",
                      "type": "const int32_t *"
                    },
                    {
                      "description": "",
                      "name": "inner",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "slices",
                      "type": "float32_t *const *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static void arm_nn_axis_scatter_f32(\n    const float32_t *packed,\n    const int32_t outer,\n    const int32_t num,\n    const int32_t *sizes,\n    const int32_t inner,\n    float32_t *const *slices\n)",
                  "source": {
                    "line": 198,
                    "path": "Include/Internal/arm_nn_axis_copy_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_axis_copy_common.h#L198"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_axis_gather_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_axis_gather_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "slices",
                      "type": "const float32_t *const *"
                    },
                    {
                      "description": "",
                      "name": "outer",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "num",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "sizes",
                      "type": "const int32_t *"
                    },
                    {
                      "description": "",
                      "name": "inner",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "packed",
                      "type": "float32_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static void arm_nn_axis_gather_f32(\n    const float32_t *const *slices,\n    const int32_t outer,\n    const int32_t num,\n    const int32_t *sizes,\n    const int32_t inner,\n    float32_t *packed\n)",
                  "source": {
                    "line": 198,
                    "path": "Include/Internal/arm_nn_axis_copy_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_axis_copy_common.h#L198"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_axis_scatter_f16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_axis_scatter_f16",
                  "params": [
                    {
                      "description": "",
                      "name": "packed",
                      "type": "const float16_t *"
                    },
                    {
                      "description": "",
                      "name": "outer",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "num",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "sizes",
                      "type": "const int32_t *"
                    },
                    {
                      "description": "",
                      "name": "inner",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "slices",
                      "type": "float16_t *const *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static void arm_nn_axis_scatter_f16(\n    const float16_t *packed,\n    const int32_t outer,\n    const int32_t num,\n    const int32_t *sizes,\n    const int32_t inner,\n    float16_t *const *slices\n)",
                  "source": {
                    "line": 201,
                    "path": "Include/Internal/arm_nn_axis_copy_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_axis_copy_common.h#L201"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_axis_gather_f16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_axis_gather_f16",
                  "params": [
                    {
                      "description": "",
                      "name": "slices",
                      "type": "const float16_t *const *"
                    },
                    {
                      "description": "",
                      "name": "outer",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "num",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "sizes",
                      "type": "const int32_t *"
                    },
                    {
                      "description": "",
                      "name": "inner",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "packed",
                      "type": "float16_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static void arm_nn_axis_gather_f16(\n    const float16_t *const *slices,\n    const int32_t outer,\n    const int32_t num,\n    const int32_t *sizes,\n    const int32_t inner,\n    float16_t *packed\n)",
                  "source": {
                    "line": 201,
                    "path": "Include/Internal/arm_nn_axis_copy_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_axis_copy_common.h#L201"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_nn_broadcast_walk.h",
              "path": "heliaCORE.Internal.arm_nn_broadcast_walk",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "Check that one NHWC dimension of two operands broadcasts to the output dimension.\n\nTensorFlow Lite broadcast rules: both inputs must be at least 1 (an empty tensor is rejected rather than treated as a no-op), they must be equal or one of them must be 1, and the output dimension must be the larger of the two.",
                  "examples": [],
                  "id": "arm_nn_broadcast_dim_valid",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_broadcast_dim_valid",
                  "params": [
                    {
                      "description": "",
                      "name": "dim_1",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "dim_2",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "dim_out",
                      "type": "const int32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static int32_t arm_nn_broadcast_dim_valid(const int32_t dim_1, const int32_t dim_2, const int32_t dim_out)",
                  "source": {
                    "line": 34,
                    "path": "Include/Internal/arm_nn_broadcast_walk.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_broadcast_walk.h#L34"
                  },
                  "summary": "Check that one NHWC dimension of two operands broadcasts to the output dimension."
                },
                {
                  "description": "Check that two NHWC operands broadcast to the given output shape.\n\nEvery kernel that uses ARM_NN_BROADCAST_WALK_NHWC must reject arguments that fail this check, since the walk indexes each input by its own dims and writes the output by the output dims.",
                  "examples": [],
                  "id": "arm_nn_broadcast_dims_valid",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_broadcast_dims_valid",
                  "params": [
                    {
                      "description": "",
                      "name": "dims_1",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "dims_2",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "dims_out",
                      "type": "const cmsis_nn_dims *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static int32_t arm_nn_broadcast_dims_valid(\n    const cmsis_nn_dims *dims_1,\n    const cmsis_nn_dims *dims_2,\n    const cmsis_nn_dims *dims_out\n)",
                  "source": {
                    "line": 54,
                    "path": "Include/Internal/arm_nn_broadcast_walk.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_broadcast_walk.h#L54"
                  },
                  "summary": "Check that two NHWC operands broadcast to the given output shape."
                },
                {
                  "description": "Walk an NHWC broadcast of two operands, calling a contiguous kernel on each run.\n\nEach input is indexed by its own dims: a dimension of 1 has stride 0 and is broadcast, any other dimension equals the output dimension and strides normally. The longest contiguous run whose shapes agree is handed to the caller's kernels, so the common cases (identical shapes, a single scalar, per-batch, per-row, per-channel) each cost one call per run.\n\nPreconditions: arm_nn_broadcast_dims_valid(dims_1, dims_2, dims_out) is non-zero.",
                  "examples": [],
                  "id": "ARM_NN_BROADCAST_WALK_NHWC",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_BROADCAST_WALK_NHWC",
                  "params": [
                    {
                      "description": "element type of the inputs",
                      "name": "IN_TYPE"
                    },
                    {
                      "description": "element type of the output",
                      "name": "OUT_TYPE"
                    },
                    {
                      "description": "const IN_TYPE * first input",
                      "name": "in_1"
                    },
                    {
                      "description": "const `cmsis_nn_dims` * dims of in_1",
                      "name": "dims_1"
                    },
                    {
                      "description": "const IN_TYPE * second input",
                      "name": "in_2"
                    },
                    {
                      "description": "const `cmsis_nn_dims` * dims of in_2",
                      "name": "dims_2"
                    },
                    {
                      "description": "OUT_TYPE * output, sized by dims_out",
                      "name": "out"
                    },
                    {
                      "description": "const `cmsis_nn_dims` * broadcast output dims (see arm_nn_broadcast_dims_valid)",
                      "name": "dims_out"
                    },
                    {
                      "description": "FULL(const IN_TYPE *a, const IN_TYPE *b, OUT_TYPE *o, int32_t n) elementwise kernel over n elements of a and b",
                      "name": "FULL"
                    },
                    {
                      "description": "SCALAR_1(const IN_TYPE *scalar, const IN_TYPE *vec, OUT_TYPE *o, int32_t n) kernel where *scalar is one element of in_1 broadcast against n elements of in_2",
                      "name": "SCALAR_1"
                    },
                    {
                      "description": "SCALAR_2(const IN_TYPE *scalar, const IN_TYPE *vec, OUT_TYPE *o, int32_t n) kernel where *scalar is one element of in_2 broadcast against n elements of in_1; note the operands arrive in reversed order, so an asymmetric kernel must swap them",
                      "name": "SCALAR_2"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_BROADCAST_WALK_NHWC(IN_TYPE, OUT_TYPE, in_1, dims_1, in_2, dims_2, out, dims_out, FULL, SCALAR_1, SCALAR_2)",
                  "source": {
                    "line": 89,
                    "path": "Include/Internal/arm_nn_broadcast_walk.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_broadcast_walk.h#L89"
                  },
                  "summary": "Walk an NHWC broadcast of two operands, calling a contiguous kernel on each run."
                }
              ]
            },
            {
              "description": "",
              "name": "arm_nn_config.h",
              "path": "heliaCORE.Internal.arm_nn_config",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "Optional feature gates for floating-point extensions.\n\nThese are disabled by default so integer-only builds keep their current code size and API surface. The float16 feature gate still depends on toolchain and target support such as ARM_FLOAT16_SUPPORTED where applicable.",
                  "examples": [],
                  "id": "ARM_NN_FLOAT_API_ENABLED",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_FLOAT_API_ENABLED",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_FLOAT_API_ENABLED (ARM_NN_ENABLE_F32 || ARM_NN_ENABLE_F16)",
                  "source": {
                    "line": 62,
                    "path": "Include/Internal/arm_nn_config.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_config.h#L62"
                  },
                  "summary": "Optional feature gates for floating-point extensions."
                }
              ]
            },
            {
              "description": "",
              "name": "arm_nn_pool_window_common.h",
              "path": "heliaCORE.Internal.arm_nn_pool_window_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_pool_axis_valid",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_pool_axis_valid",
                  "params": [
                    {
                      "description": "",
                      "name": "n",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "s",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "p",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "k",
                      "type": "const int32_t"
                    },
                    {
                      "description": "",
                      "name": "w",
                      "type": "const int32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_nn_pool_axis_valid(const int32_t n, const int32_t s, const int32_t p, const int32_t k, const int32_t w)",
                  "source": {
                    "line": 40,
                    "path": "Include/Internal/arm_nn_pool_window_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_pool_window_common.h#L40"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_nn_s4_decode.h",
              "path": "heliaCORE.Internal.arm_nn_s4_decode",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_s4_low_nibble",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_s4_low_nibble",
                  "params": [
                    {
                      "description": "",
                      "name": "packed",
                      "type": "int8_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static int8_t arm_nn_s4_low_nibble(int8_t packed)",
                  "source": {
                    "line": 17,
                    "path": "Include/Internal/arm_nn_s4_decode.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_s4_decode.h#L17"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_s4_high_nibble",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_s4_high_nibble",
                  "params": [
                    {
                      "description": "",
                      "name": "packed",
                      "type": "int8_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static int8_t arm_nn_s4_high_nibble(int8_t packed)",
                  "source": {
                    "line": 25,
                    "path": "Include/Internal/arm_nn_s4_decode.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_s4_decode.h#L25"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_nn_sqrt_flt.h",
              "path": "heliaCORE.Internal.arm_nn_sqrt_flt",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_NN_SQRT_EXACT_FN",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_NN_SQRT_EXACT_FN",
                  "params": [],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_NN_SQRT_EXACT_FN",
                  "source": {
                    "line": 53,
                    "path": "Include/Internal/arm_nn_sqrt_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_sqrt_flt.h#L53"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_f32_to_bits",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_f32_to_bits",
                  "params": [
                    {
                      "description": "",
                      "name": "value",
                      "type": "float32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static uint32_t arm_nn_f32_to_bits(float32_t value)",
                  "source": {
                    "line": 58,
                    "path": "Include/Internal/arm_nn_sqrt_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_sqrt_flt.h#L58"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_f32_from_bits",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_f32_from_bits",
                  "params": [
                    {
                      "description": "",
                      "name": "bits",
                      "type": "uint32_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float32_t arm_nn_f32_from_bits(uint32_t bits)",
                  "source": {
                    "line": 65,
                    "path": "Include/Internal/arm_nn_sqrt_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_sqrt_flt.h#L65"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_sqrt_special_f32",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_sqrt_special_f32",
                  "params": [
                    {
                      "description": "",
                      "name": "in_bits",
                      "type": "uint32_t"
                    },
                    {
                      "description": "",
                      "name": "reciprocal",
                      "type": "bool"
                    },
                    {
                      "description": "",
                      "name": "out_bits",
                      "type": "uint32_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_nn_sqrt_special_f32(uint32_t in_bits, bool reciprocal, uint32_t *out_bits)",
                  "source": {
                    "line": 73,
                    "path": "Include/Internal/arm_nn_sqrt_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_sqrt_flt.h#L73"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_f16_to_bits",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_f16_to_bits",
                  "params": [
                    {
                      "description": "",
                      "name": "value",
                      "type": "float16_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static uint16_t arm_nn_f16_to_bits(float16_t value)",
                  "source": {
                    "line": 104,
                    "path": "Include/Internal/arm_nn_sqrt_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_sqrt_flt.h#L104"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_f16_from_bits",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_f16_from_bits",
                  "params": [
                    {
                      "description": "",
                      "name": "bits",
                      "type": "uint16_t"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static float16_t arm_nn_f16_from_bits(uint16_t bits)",
                  "source": {
                    "line": 111,
                    "path": "Include/Internal/arm_nn_sqrt_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_sqrt_flt.h#L111"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_sqrt_special_f16",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_sqrt_special_f16",
                  "params": [
                    {
                      "description": "",
                      "name": "in_bits",
                      "type": "uint16_t"
                    },
                    {
                      "description": "",
                      "name": "reciprocal",
                      "type": "bool"
                    },
                    {
                      "description": "",
                      "name": "out_bits",
                      "type": "uint16_t *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static bool arm_nn_sqrt_special_f16(uint16_t in_bits, bool reciprocal, uint16_t *out_bits)",
                  "source": {
                    "line": 119,
                    "path": "Include/Internal/arm_nn_sqrt_flt.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_nn_sqrt_flt.h#L119"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_resize_nearest_neighbor_common.h",
              "path": "heliaCORE.Internal.arm_resize_nearest_neighbor_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_resize_nearest_neighbor_scratch_bytes",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_resize_nearest_neighbor_scratch_bytes",
                  "params": [
                    {
                      "description": "",
                      "name": "output_dims",
                      "type": "const cmsis_nn_dims *"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static int32_t arm_nn_resize_nearest_neighbor_scratch_bytes(const cmsis_nn_dims *output_dims)",
                  "source": {
                    "line": 39,
                    "path": "Include/Internal/arm_resize_nearest_neighbor_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_resize_nearest_neighbor_common.h#L39"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "arm_nn_resize_nearest_neighbor_prepare",
                  "kind": "function",
                  "language": "c",
                  "members": [],
                  "name": "arm_nn_resize_nearest_neighbor_prepare",
                  "params": [
                    {
                      "description": "",
                      "name": "ctx",
                      "type": "const cmsis_nn_context *"
                    },
                    {
                      "description": "",
                      "name": "resize_params",
                      "type": "const cmsis_nn_resize_params *"
                    },
                    {
                      "description": "",
                      "name": "input_shape",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_size_shape",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "output_size_data",
                      "type": "const int32_t *"
                    },
                    {
                      "description": "",
                      "name": "output_shape",
                      "type": "const cmsis_nn_dims *"
                    },
                    {
                      "description": "",
                      "name": "x_map_out",
                      "type": "int32_t **"
                    },
                    {
                      "description": "",
                      "name": "y_map_out",
                      "type": "int32_t **"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "static arm_cmsis_nn_status arm_nn_resize_nearest_neighbor_prepare(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_resize_params *resize_params,\n    const cmsis_nn_dims *input_shape,\n    const cmsis_nn_dims *output_size_shape,\n    const int32_t *output_size_data,\n    const cmsis_nn_dims *output_shape,\n    int32_t **x_map_out,\n    int32_t **y_map_out\n)",
                  "source": {
                    "line": 50,
                    "path": "Include/Internal/arm_resize_nearest_neighbor_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_resize_nearest_neighbor_common.h#L50"
                  },
                  "summary": ""
                },
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_RESIZE_NEAREST_NEIGHBOR_DEFINE",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_RESIZE_NEAREST_NEIGHBOR_DEFINE",
                  "params": [
                    {
                      "description": "",
                      "name": "FUNC_NAME"
                    },
                    {
                      "description": "",
                      "name": "SCALAR_T"
                    },
                    {
                      "description": "",
                      "name": "MEMCPY_FUNC"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_RESIZE_NEAREST_NEIGHBOR_DEFINE(FUNC_NAME, SCALAR_T, MEMCPY_FUNC)",
                  "source": {
                    "line": 129,
                    "path": "Include/Internal/arm_resize_nearest_neighbor_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_resize_nearest_neighbor_common.h#L129"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_strided_slice_common.h",
              "path": "heliaCORE.Internal.arm_strided_slice_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_STRIDED_SLICE_DEFINE",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_STRIDED_SLICE_DEFINE",
                  "params": [
                    {
                      "description": "",
                      "name": "FUNC_NAME"
                    },
                    {
                      "description": "",
                      "name": "SCALAR_T"
                    },
                    {
                      "description": "",
                      "name": "MEMCPY_FUNC"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_STRIDED_SLICE_DEFINE(FUNC_NAME, SCALAR_T, MEMCPY_FUNC)",
                  "source": {
                    "line": 32,
                    "path": "Include/Internal/arm_strided_slice_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_strided_slice_common.h#L32"
                  },
                  "summary": ""
                }
              ]
            },
            {
              "description": "",
              "name": "arm_transpose_common.h",
              "path": "heliaCORE.Internal.arm_transpose_common",
              "submodules": [],
              "summary": "",
              "symbols": [
                {
                  "description": "",
                  "examples": [],
                  "id": "ARM_TRANSPOSE_DEFINE",
                  "kind": "macro",
                  "language": "c",
                  "members": [],
                  "name": "ARM_TRANSPOSE_DEFINE",
                  "params": [
                    {
                      "description": "",
                      "name": "FUNC_NAME"
                    },
                    {
                      "description": "",
                      "name": "SCALAR_T"
                    },
                    {
                      "description": "",
                      "name": "PARAMS_T"
                    },
                    {
                      "description": "",
                      "name": "LAYOUT_T"
                    },
                    {
                      "description": "",
                      "name": "TRANSPOSE_2D_FUNC"
                    }
                  ],
                  "raises": [],
                  "returns": [],
                  "signature": "#define ARM_TRANSPOSE_DEFINE(FUNC_NAME, SCALAR_T, PARAMS_T, LAYOUT_T, TRANSPOSE_2D_FUNC)",
                  "source": {
                    "line": 35,
                    "path": "Include/Internal/arm_transpose_common.h",
                    "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/Internal/arm_transpose_common.h#L35"
                  },
                  "summary": ""
                }
              ]
            }
          ],
          "summary": "",
          "symbols": []
        },
        {
          "description": "",
          "name": "arm_nn_math_types_flt.h",
          "path": "heliaCORE.arm_nn_math_types_flt",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "32-bit floating-point type definition for CMSIS-NN float extensions.",
              "examples": [],
              "id": "float32_t",
              "kind": "type",
              "language": "c",
              "members": [],
              "name": "float32_t",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "typedef float float32_t",
              "source": {
                "line": 45,
                "path": "Include/arm_nn_math_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types_flt.h#L45"
              },
              "summary": "32-bit floating-point type definition for CMSIS-NN float extensions."
            },
            {
              "description": "Largest finite float32 value representable by the toolchain.",
              "examples": [],
              "id": "ARM_NN_F32_FINITE_MAX",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_F32_FINITE_MAX",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_F32_FINITE_MAX ((float32_t)__FLT_MAX__)",
              "source": {
                "line": 52,
                "path": "Include/arm_nn_math_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types_flt.h#L52"
              },
              "summary": "Largest finite float32 value representable by the toolchain."
            },
            {
              "description": "Lowest finite float32 value representable by the toolchain.",
              "examples": [],
              "id": "ARM_NN_F32_FINITE_LOWEST",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_F32_FINITE_LOWEST",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_F32_FINITE_LOWEST (-ARM_NN_F32_FINITE_MAX)",
              "source": {
                "line": 57,
                "path": "Include/arm_nn_math_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types_flt.h#L57"
              },
              "summary": "Lowest finite float32 value representable by the toolchain."
            },
            {
              "description": "Largest finite float16 value representable by the toolchain.",
              "examples": [],
              "id": "ARM_NN_F16_FINITE_MAX",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_F16_FINITE_MAX",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_F16_FINITE_MAX ((float16_t)65504.0f)",
              "source": {
                "line": 116,
                "path": "Include/arm_nn_math_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types_flt.h#L116"
              },
              "summary": "Largest finite float16 value representable by the toolchain."
            },
            {
              "description": "Lowest finite float16 value representable by the toolchain.\n\nNegated before the cast: negating a float16_t that is __fp16 promotes it to float, which -Wdouble-promotion reports at every use.",
              "examples": [],
              "id": "ARM_NN_F16_FINITE_LOWEST",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_F16_FINITE_LOWEST",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_F16_FINITE_LOWEST ((float16_t)(-65504.0f))",
              "source": {
                "line": 128,
                "path": "Include/arm_nn_math_types_flt.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types_flt.h#L128"
              },
              "summary": "Lowest finite float16 value representable by the toolchain."
            }
          ]
        },
        {
          "description": "",
          "name": "arm_nn_math_types.h",
          "path": "heliaCORE.arm_nn_math_types",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "Translate architecture feature flags to CMSIS-NN defines.\n\nLimits macros",
              "examples": [],
              "id": "NN_Q31_MAX",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NN_Q31_MAX",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NN_Q31_MAX ((int32_t)(0x7FFFFFFFL))",
              "source": {
                "line": 117,
                "path": "Include/arm_nn_math_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types.h#L117"
              },
              "summary": "Translate architecture feature flags to CMSIS-NN defines."
            },
            {
              "description": "",
              "examples": [],
              "id": "NN_Q15_MAX",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NN_Q15_MAX",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NN_Q15_MAX ((int16_t)(0x7FFF))",
              "source": {
                "line": 118,
                "path": "Include/arm_nn_math_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types.h#L118"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "NN_Q7_MAX",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NN_Q7_MAX",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NN_Q7_MAX ((int8_t)(0x7F))",
              "source": {
                "line": 119,
                "path": "Include/arm_nn_math_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types.h#L119"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "NN_Q31_MIN",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NN_Q31_MIN",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NN_Q31_MIN ((int32_t)(0x80000000L))",
              "source": {
                "line": 120,
                "path": "Include/arm_nn_math_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types.h#L120"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "NN_Q15_MIN",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NN_Q15_MIN",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NN_Q15_MIN ((int16_t)(0x8000))",
              "source": {
                "line": 121,
                "path": "Include/arm_nn_math_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types.h#L121"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "NN_Q7_MIN",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NN_Q7_MIN",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NN_Q7_MIN ((int8_t)(0x80))",
              "source": {
                "line": 122,
                "path": "Include/arm_nn_math_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_math_types.h#L122"
              },
              "summary": ""
            }
          ]
        },
        {
          "description": "",
          "name": "arm_nn_tables.h",
          "path": "heliaCORE.arm_nn_tables",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "tables for various activation functions",
              "examples": [],
              "id": "sigmoid_table_uint16",
              "kind": "attribute",
              "language": "c",
              "members": [],
              "name": "sigmoid_table_uint16",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "const uint16_t sigmoid_table_uint16[256]",
              "source": {
                "line": 40,
                "path": "Include/arm_nn_tables.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_tables.h#L40"
              },
              "summary": "tables for various activation functions"
            }
          ]
        },
        {
          "description": "",
          "name": "arm_nn_types.h",
          "path": "heliaCORE.arm_nn_types",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "",
              "examples": [],
              "id": "NS_CMSIS_NN_VERSION_MAJOR",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NS_CMSIS_NN_VERSION_MAJOR",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NS_CMSIS_NN_VERSION_MAJOR (7) /* x-release-please-major */",
              "source": {
                "line": 41,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L41"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "NS_CMSIS_NN_VERSION_MINOR",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NS_CMSIS_NN_VERSION_MINOR",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NS_CMSIS_NN_VERSION_MINOR (39) /* x-release-please-minor */",
              "source": {
                "line": 42,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L42"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "NS_CMSIS_NN_VERSION_PATCH",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NS_CMSIS_NN_VERSION_PATCH",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NS_CMSIS_NN_VERSION_PATCH (2) /* x-release-please-patch */",
              "source": {
                "line": 43,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L43"
              },
              "summary": ""
            },
            {
              "description": "Identity macros for the ns-cmsis-nn (Ambiq) superset of CMSIS-NN.\n\nThis library is wire-compatible with upstream ARM-software/CMSIS-NN: every upstream `arm_*` symbol resolves here. We additionally ship Ambiq-specific kernels (e.g. arm_gather_s8, the elementwise prelu/clamp variants, ...).\n\nDownstream code that depends on Ambiq-only kernels should guard with:\n\n```\n#if !defined(NS_CMSIS_NN)\n#  error \"this code requires ns-cmsis-nn (Ambiq superset)\"\n#endif\n#if NS_CMSIS_NN_VERSION < 7024000\n#  error \"needs ns-cmsis-nn >= 7.24.0\"\n#endif\n```\n\nNS_CMSIS_NN_VERSION is packed as MAJOR * 1000000 + MINOR * 1000 + PATCH (each of MINOR and PATCH gets a full 3-digit field, so semantic ordering is preserved for any reasonable version). It tracks release-please bumps automatically through the per-component macros above; no separate marker is required.",
              "examples": [],
              "id": "NS_CMSIS_NN",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NS_CMSIS_NN",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NS_CMSIS_NN (1)",
              "source": {
                "line": 67,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L67"
              },
              "summary": "Identity macros for the ns-cmsis-nn (Ambiq) superset of CMSIS-NN."
            },
            {
              "description": "",
              "examples": [],
              "id": "NS_CMSIS_NN_VERSION",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NS_CMSIS_NN_VERSION",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define NS_CMSIS_NN_VERSION ((NS_CMSIS_NN_VERSION_MAJOR * 1000000) + (NS_CMSIS_NN_VERSION_MINOR * 1000) + NS_CMSIS_NN_VERSION_PATCH)",
              "source": {
                "line": 68,
                "path": "Include/arm_nn_types.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nn_types.h#L68"
              },
              "summary": ""
            }
          ]
        },
        {
          "description": "",
          "name": "arm_nnfunctions.h",
          "path": "heliaCORE.arm_nnfunctions",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "",
              "examples": [],
              "id": "USE_INTRINSIC",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "USE_INTRINSIC",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define USE_INTRINSIC",
              "source": {
                "line": 47,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L47"
              },
              "summary": ""
            },
            {
              "description": "s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.",
              "examples": [],
              "id": "arm_convolve_wrapper_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s4",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_wrapper_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8 packed with 2x int4",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_wrapper_s4(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 97,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L97"
              },
              "summary": "s4 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution."
            },
            {
              "description": "Get the required buffer size for arm_convolve_wrapper_s4.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.",
              "examples": [],
              "id": "arm_convolve_wrapper_s4_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s4_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_s4_get_buffer_size(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 133,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L133"
              },
              "summary": "Get the required buffer size for armconvolvewrappers4."
            },
            {
              "description": "Get the required buffer size for arm_convolve_wrapper_s4 for Arm(R) Helium Architecture case.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s4_get_buffer_size()`. Currently this operator does not have an mve implementation, so dsp will be used.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a 0 from a route needing no buffer is not a statement that the shape is valid.\n\n:::",
              "examples": [],
              "id": "arm_convolve_wrapper_s4_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s4_get_buffer_size_mve",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_s4_get_buffer_size_mve(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 149,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L149"
              },
              "summary": "Get the required buffer size for armconvolvewrappers4 for Arm(R) Helium Architecture case."
            },
            {
              "description": "Get the required buffer size for arm_convolve_wrapper_s4 for processors with DSP extension.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size. Which routes compute a byte count is build-dependent - the 1x1 routes need no buffer on any build, though the 1x1 fast route still rejects a negative input_dims->c with -1, and on a Helium build the 1xN route needs none when its padding lines up with the stride - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s4_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a 0 from a route needing no buffer is not a statement that the shape is valid.\n\n:::",
              "examples": [],
              "id": "arm_convolve_wrapper_s4_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s4_get_buffer_size_dsp",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_s4_get_buffer_size_dsp(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 164,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L164"
              },
              "summary": "Get the required buffer size for armconvolvewrappers4 for processors with DSP extension."
            },
            {
              "description": "s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.\n\n- On builds with ARM_MATH_MVEI (without ARM_MATH_AUTOVECTORIZE), a layer that would run `arm_convolve_s8()` and is in the gate of `arm_convolve_s8_small_cin()` or `arm_convolve_s8_3x3_c16_s1()` runs that entry instead, with the same result, scratch and weight sums. The input depth is checked first, so other layers skip both gates.",
              "examples": [],
              "id": "arm_convolve_wrapper_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s8",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_wrapper_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-output-channel weight sums, supplied by the caller. The selected kernel only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build: this wrapper dispatches to `arm_convolve_s8()`, `arm_convolve_1x1_s8()`, `arm_convolve_1x1_s8_fast()`, `arm_convolve_1_x_n_s8()` and `arm_convolve_1x1_out_s8()`. The buffer contents are consumed only on builds with the MVE extension (ARM_MATH_MVEI), and on those builds every one of those kernels diagnoses a NULL buf with ARM_CMSIS_NN_ARG_ERROR; on other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed that way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_wrapper_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 225,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L225"
              },
              "summary": "s8 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution."
            },
            {
              "description": "Get the required buffer size for arm_convolve_wrapper_s8.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.",
              "examples": [],
              "id": "arm_convolve_wrapper_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s8_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_s8_get_buffer_size(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 264,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L264"
              },
              "summary": "Get the required buffer size for armconvolvewrappers8."
            },
            {
              "description": "Get the required buffer size for arm_convolve_s8 for Arm(R) Helium Architecture case.\n\nThe dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas, so a shape whose dimension product overflows is only reported by the leg that actually computes a buffer for it.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_convolve_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher.\n\n:::",
              "examples": [],
              "id": "arm_convolve_s8_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s8_get_buffer_size_mve",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_convolve_s8_get_buffer_size_mve(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 278,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L278"
              },
              "summary": "Get the required buffer size for armconvolves8 for Arm(R) Helium Architecture case."
            },
            {
              "description": "Get the required buffer size for arm_convolve_wrapper_s8 for Arm(R) Helium Architecture case.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a 0 from a route needing no buffer is not a statement that the shape is valid.\n\n:::",
              "examples": [],
              "id": "arm_convolve_wrapper_s8_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s8_get_buffer_size_mve",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_s8_get_buffer_size_mve(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 290,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L290"
              },
              "summary": "Get the required buffer size for armconvolvewrappers8 for Arm(R) Helium Architecture case."
            },
            {
              "description": "Get the required buffer size for arm_convolve_wrapper_s8 for processors with DSP extension.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and where this function composes sub-sizer results the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent, and on a DSP build also compiler-dependent - the 1x1 route that is not the fast variant needs no buffer on any build, and the 1x1 fast route needs none on a DSP build outside armclang - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a 0 from a route needing no buffer is not a statement that the shape is valid.\n\n:::",
              "examples": [],
              "id": "arm_convolve_wrapper_s8_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s8_get_buffer_size_dsp",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. Which dimensions a route inspects is build-dependent, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_s8_get_buffer_size_dsp(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 305,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L305"
              },
              "summary": "Get the required buffer size for armconvolvewrappers8 for processors with DSP extension."
            },
            {
              "description": "s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution.",
              "examples": [],
              "id": "arm_convolve_wrapper_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s16",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_wrapper_s8_get_buffer_size will return the buffer_size if required The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const cmsis_nn_bias_data *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_wrapper_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const cmsis_nn_bias_data *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 338,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L338"
              },
              "summary": "s16 convolution layer wrapper function with the main purpose to call the optimal kernel available in cmsis-nn to perform the convolution."
            },
            {
              "description": "s16 grouped convolution optimized for the case where filter_dims->c == 1 and input_ch == output_ch (channel multiplier = 1).",
              "examples": [],
              "id": "arm_convolve_s16_group_ch_mult_1",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s16_group_ch_mult_1",
              "params": [
                {
                  "description": "Function context (unused, pass NULL-initialised).",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (strides, dilations, pads, activation).",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info (multiplier and shift).",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, 1]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions (unused, may be zero-initialised).",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias struct (int32 or int64). May be NULL.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const cmsis_nn_bias_data *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_s16_group_ch_mult_1(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const cmsis_nn_bias_data *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 368,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L368"
              },
              "summary": "s16 grouped convolution optimized for the case where filterdims->c == 1 and inputch == outputch (channel multiplier = 1)."
            },
            {
              "description": "Get the required buffer size for arm_convolve_wrapper_s16.\n\nAn out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size.",
              "examples": [],
              "id": "arm_convolve_wrapper_s16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s16_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_s16_get_buffer_size(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 398,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L398"
              },
              "summary": "Get the required buffer size for armconvolvewrappers16."
            },
            {
              "description": "Get the required buffer size for arm_convolve_wrapper_s16 for for processors with DSP extension.\n\nAn out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s16_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher.\n\n:::",
              "examples": [],
              "id": "arm_convolve_wrapper_s16_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s16_get_buffer_size_dsp",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_s16_get_buffer_size_dsp(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 412,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L412"
              },
              "summary": "Get the required buffer size for armconvolvewrappers16 for for processors with DSP extension."
            },
            {
              "description": "Get the required buffer size for arm_convolve_wrapper_s16 for Arm(R) Helium Architecture case.\n\nAn out-of-range shape is reported as -1 rather than a wrapped size. Where this function composes sub-sizer results, the sentinel is propagated before any `ARM_NN_MAX()` or sum, so it can never collapse into a plausible positive size.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_convolve_wrapper_s16_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher.\n\n:::",
              "examples": [],
              "id": "arm_convolve_wrapper_s16_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_wrapper_s16_get_buffer_size_mve",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_convolve_wrapper_s16_get_buffer_size_mve(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 426,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L426"
              },
              "summary": "Get the required buffer size for armconvolvewrappers16 for Arm(R) Helium Architecture case."
            },
            {
              "description": "Basic s4 convolution function.\n\n1. Supported framework: TensorFlow Lite micro\n2. Additional memory is required for optimization. Refer to argument 'ctx' for details.",
              "examples": [],
              "id": "arm_convolve_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s4",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Packed Filter data pointer. Data type: int8 packed with 2x int4",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_s4(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 458,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L458"
              },
              "summary": "Basic s4 convolution function."
            },
            {
              "description": "Basic s4 convolution function with a requirement of even number of kernels.\n\n1. Supported framework: TensorFlow Lite micro\n2. Additional memory is required for optimization. Refer to argument 'ctx' for details.",
              "examples": [],
              "id": "arm_convolve_even_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_even_s4",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_even_s4_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions. Note the product must be even.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Packed Filter data pointer. Data type: int8 packed with 2x int4",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR` if not for MVE"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_even_s4(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 499,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L499"
              },
              "summary": "Basic s4 convolution function with a requirement of even number of kernels."
            },
            {
              "description": "Basic s8 convolution function.\n\n1. Supported framework: TensorFlow Lite micro\n2. Additional memory is required for optimization. Refer to argument 'ctx' for details.",
              "examples": [],
              "id": "arm_convolve_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s8",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. For grouped convolution the entries run over all output_dims->c channels, groups laid out consecutively. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI); an unfilled buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, and a NULL buf is currently reported as ARM_CMSIS_NN_ARG_ERROR. On other builds this function currently derives the same quantity itself and does not read the context. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, CK] where HK, WK and CK are the spatial filter dimensions. CK != C_IN is used for grouped convolution, in which case the required conditions are C_IN = N * CK and C_OUT = N * M for N groups of size M.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Upscale tensor dimensions for transpose. Format: [H_UP, W_UP]",
                  "direction": "in",
                  "name": "upscale_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *upscale_dims,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 563,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L563"
              },
              "summary": "Basic s8 convolution function."
            },
            {
              "description": "s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model. It copies each kernel row with one predicated vector load and multiplies four output channels per step.\n\n- The output is identical to `arm_convolve_s8()`. The bias is read through the weight sums, which `arm_convolve_weight_sum()` fills as for `arm_convolve_s8()`; bias_dims and bias_data are unused.\n- Gate: upscale_dims NULL, C_IN from 1 to 3 with CK equal to C_IN (one group), dilation 1 in both dimensions, WK and HK at least 1 with WK x C_IN at most 16 and HK x WK x C_IN at most 48, and C_OUT a positive multiple of 4. Stride, padding and batch count are as for `arm_convolve_s8()`.\n- Scratch: ctx->buf holds `arm_convolve_s8_get_buffer_size()` bytes (4 x 16 x ceil(HK x WK x C_IN / 16) on ARM_MATH_MVEI builds), the same as `arm_convolve_s8()`, and needs no alignment.\n- It is a direct entry: `arm_convolve_s8()` does not call it, and `arm_convolve_wrapper_s8()` calls it for layers in the gate that it would otherwise pass to `arm_convolve_s8()`. A caller that selects the kernel per layer ahead of time calls it for layers in the gate and `arm_convolve_s8()` for every other layer, or on `ARM_CMSIS_NN_NO_IMPL_ERROR`. Both take the same arguments, scratch and weight sums.",
              "examples": [],
              "id": "arm_convolve_s8_small_cin",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s8_small_cin",
              "params": [
                {
                  "description": "Function context with `arm_convolve_s8_get_buffer_size()` bytes of scratch, all of which may be written",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-output-channel weight sums, as for `arm_convolve_s8()`",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters, as for `arm_convolve_s8()`",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, CK]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Upscale tensor dimensions for transpose. Format: [H_UP, W_UP]",
                  "direction": "in",
                  "name": "upscale_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - an argument error that `arm_convolve_s8()` reports: ctx->buf is NULL, C_IN or C_OUT is not a multiple of the group count C_IN / CK, or weight_sum_ctx->buf is NULL on builds with ARM_MATH_MVEI. These are checked before the gate. `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer is outside the gate below, or the build lacks ARM_MATH_MVEI or defines ARM_MATH_AUTOVECTORIZE; nothing is written, to the output or to the scratch `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_s8_small_cin(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *upscale_dims,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 620,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L620"
              },
              "summary": "s8 convolution for input depths of 1 to 3, such as the first layer of an image or audio model."
            },
            {
              "description": "s8 3x3 convolution over 16 input channels with unit stride. It reads the kernel rows of a patch inside the input in place, copying only patches that cross the border, and multiplies four output pixels per filter load.\n\n- The output is identical to `arm_convolve_s8()`. The bias is read through the weight sums; bias_dims and bias_data are unused.\n- Gate: upscale_dims NULL, C_IN and CK both 16 (one group), HK and WK both 3, and stride and dilation 1 in both dimensions. Padding, batch count and C_OUT are as for `arm_convolve_s8()`.\n- It is a direct entry: `arm_convolve_s8()` does not call it, and `arm_convolve_wrapper_s8()` calls it for layers in the gate that it would otherwise pass to `arm_convolve_s8()`. A caller that selects the kernel per layer ahead of time calls it for layers in the gate and `arm_convolve_s8()` for every other layer, or on `ARM_CMSIS_NN_NO_IMPL_ERROR`. The gate does not overlap that of `arm_convolve_s8_small_cin()`.",
              "examples": [],
              "id": "arm_convolve_s8_3x3_c16_s1",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s8_3x3_c16_s1",
              "params": [
                {
                  "description": "Function context with `arm_convolve_s8_get_buffer_size()` bytes of scratch (576 bytes on ARM_MATH_MVEI builds), all of which may be written",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-output-channel weight sums, as for `arm_convolve_s8()`",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters, as for `arm_convolve_s8()`",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, CK]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Upscale tensor dimensions for transpose. Format: [H_UP, W_UP]",
                  "direction": "in",
                  "name": "upscale_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_convolve_s8_small_cin()` `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer is outside the gate below, or the build lacks ARM_MATH_MVEI or defines ARM_MATH_AUTOVECTORIZE; nothing is written, to the output or to the scratch `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_s8_3x3_c16_s1(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *upscale_dims,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 672,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L672"
              },
              "summary": "s8 3x3 convolution over 16 input channels with unit stride."
            },
            {
              "description": "Get the required buffer size for s4 convolution function.\n\nThe dimensions and the byte count are both checked here, so an out-of-range shape returns -1 on every build target rather than a wrapped size.",
              "examples": [],
              "id": "arm_convolve_s4_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s4_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_convolve_s4_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 698,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L698"
              },
              "summary": "Get the required buffer size for s4 convolution function."
            },
            {
              "description": "Get the required buffer size for arm_convolve_even_s4.\n\nForwards to `arm_convolve_s4_get_buffer_size()`: the even_s4 kernel stages up to four im2col rows of filter_dims->w * filter_dims->h * input_dims->c int8 elements, byte-for-byte the size that sizer returns. The equality, including the -1 answers for out-of-range shapes, is pinned by a Unity test.",
              "examples": [],
              "id": "arm_convolve_even_s4_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_even_s4_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_convolve_even_s4_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 713,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L713"
              },
              "summary": "Get the required buffer size for armconvolveevens4."
            },
            {
              "description": "Get the required buffer size for s8 convolution function.\n\nThe dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas, so a shape whose dimension product overflows is only reported by the leg that actually computes a buffer for it.",
              "examples": [],
              "id": "arm_convolve_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s8_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_convolve_s8_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 729,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L729"
              },
              "summary": "Get the required buffer size for s8 convolution function."
            },
            {
              "description": "Get the required buffer size for s8 convolution and depthwise convolution weight sum.\n\nFor a valid (non-negative, in-range) output_dims->c, returns output_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. A negative or out-of-range output_dims->c returns -1 on builds with the MVE extension; elsewhere no weight sum buffer is used and the answer stays 0.",
              "examples": [],
              "id": "arm_convolve_s8_get_weights_sum_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s8_get_weights_sum_size",
              "params": [
                {
                  "description": "Output (activation) tensor dimensions. Format: [N, H, W, C_COUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required weight sum buffer size in bytes, or -1 if output_dims->c is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_convolve_s8_get_weights_sum_size(const cmsis_nn_dims *output_dims)",
              "source": {
                "line": 742,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L742"
              },
              "summary": "Get the required buffer size for s8 convolution and depthwise convolution weight sum."
            },
            {
              "description": "Wrapper to select optimal transposed convolution algorithm depending on parameters.\n\n1. Supported framework: TensorFlow Lite micro\n2. Additional memory is required for optimization. Refer to arguments 'ctx' and 'reverse_conv_ctx' for details.\n3. Dilation is not supported: transpose_conv_params->dilation must be 1 in both dimensions, otherwise `ARM_CMSIS_NN_ARG_ERROR` is returned.",
              "examples": [],
              "id": "arm_transpose_conv_wrapper_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_wrapper_s8",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. `arm_transpose_conv_s8_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-output-channel weight sums, supplied by the caller. The function only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and transpose_conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing transpose_conv_params->input_offset as lhs_offset and the same bias_data given here, so that entry j holds input_offset * sum(weights of output channel j) + bias_data[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Compute the sums over filter_data exactly as passed to this function: this wrapper guarantees that whatever filter preparation it performs internally preserves the per-output-channel sums, so no reversed or otherwise rearranged copy of the weights is needed for this step. Pass a valid context on every build. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the reverse-convolution route forwards this context to `arm_convolve_s8()`; an unfilled buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, and a NULL buf is currently reported as ARM_CMSIS_NN_ARG_ERROR. On other builds the contents are currently not read. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Function context for the reversed filter used when this wrapper routes to the reverse convolution. Holds filter height * filter width * input channels * output channels int8 values; `arm_transpose_conv_s8_get_reverse_conv_buffer_size()` returns the required size (0 when the reverse-convolution route is not taken). The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "reverse_conv_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of transpose_conv_params->input_offset : [-127, 128] Range of transpose_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each out channel.",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_conv_wrapper_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_context *reverse_conv_ctx,\n    const cmsis_nn_transpose_conv_params *transpose_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 812,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L812"
              },
              "summary": "Wrapper to select optimal transposed convolution algorithm depending on parameters."
            },
            {
              "description": "Basic s8 transpose convolution function.\n\n1. Supported framework: TensorFlow Lite micro\n2. Additional memory is required for optimization. Refer to argument 'ctx' for details; 'output_ctx' is unused.\n3. Dilation is not supported: transpose_conv_params->dilation must be 1 in both dimensions, otherwise `ARM_CMSIS_NN_ARG_ERROR` is returned.",
              "examples": [],
              "id": "arm_transpose_conv_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_s8",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_transpose_conv_s8_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Not accessed by this function: its buffer is neither read nor written, and it therefore has no size requirement. The parameter exists only to keep one signature across the transpose-conv family, whose float twins ignore it the same way; `arm_transpose_conv_wrapper_s8()` forwards its reverse_conv_ctx into this slot. In-tree callers pass a valid context, whose buf may be NULL.",
                  "direction": "inout",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of transpose_conv_params->input_offset : [-127, 128] Range of transpose_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "transpose_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each out channel.",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_conv_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params *transpose_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 864,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L864"
              },
              "summary": "Basic s8 transpose convolution function."
            },
            {
              "description": "Get the required buffer size for ctx in s8 transpose conv function.\n\nThe returned size is safe for both `arm_transpose_conv_s8()` and `arm_transpose_conv_wrapper_s8()`: it is the larger of the two routes' requirements, so it may exceed what the wrapper's reverse-convolution route alone would need. When either route is out of range the sentinel is propagated ahead of that comparison, so -1 is never collapsed into a plausible positive size by the other route.",
              "examples": [],
              "id": "arm_transpose_conv_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_s8_get_buffer_size",
              "params": [
                {
                  "description": "Transposed convolution parameters",
                  "direction": "in",
                  "name": "transposed_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "out_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative, either stride is not positive, or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_transpose_conv_s8_get_buffer_size(\n    const cmsis_nn_transpose_conv_params *transposed_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *out_dims\n)",
              "source": {
                "line": 894,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L894"
              },
              "summary": "Get the required buffer size for ctx in s8 transpose conv function."
            },
            {
              "description": "Get the required buffer size for output_ctx in s8 transpose conv function.",
              "examples": [],
              "id": "arm_transpose_conv_s8_get_reverse_conv_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_s8_get_reverse_conv_buffer_size",
              "params": [
                {
                  "description": "Transposed convolution parameters",
                  "direction": "in",
                  "name": "transposed_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_transpose_conv_s8_get_reverse_conv_buffer_size(\n    const cmsis_nn_transpose_conv_params *transposed_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims\n)",
              "source": {
                "line": 910,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L910"
              },
              "summary": "Get the required buffer size for outputctx in s8 transpose conv function."
            },
            {
              "description": "Get size of additional buffer required by `arm_transpose_conv_s8()` for Arm(R) Helium Architecture case.\n\nThe returned size is safe for both `arm_transpose_conv_s8()` and `arm_transpose_conv_wrapper_s8()`: it is the larger of the two routes' requirements, so it may exceed what the wrapper's reverse-convolution route alone would need. When either route is out of range the sentinel is propagated ahead of that comparison, so -1 is never collapsed into a plausible positive size by the other route.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_transpose_conv_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher.\n\n:::",
              "examples": [],
              "id": "arm_transpose_conv_s8_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_conv_s8_get_buffer_size_mve",
              "params": [
                {
                  "description": "Transposed convolution parameters",
                  "direction": "in",
                  "name": "transposed_conv_params",
                  "type": "const cmsis_nn_transpose_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "out_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative, either stride is not positive, or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_transpose_conv_s8_get_buffer_size_mve(\n    const cmsis_nn_transpose_conv_params *transposed_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *out_dims\n)",
              "source": {
                "line": 923,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L923"
              },
              "summary": "Get size of additional buffer required by armtransposeconvs8() for Arm(R) Helium Architecture case."
            },
            {
              "description": "Basic s16 convolution function.\n\n1. Supported framework: TensorFlow Lite micro\n2. Additional memory is required for optimization. Refer to argument 'ctx' for details.",
              "examples": [],
              "id": "arm_convolve_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s16",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const cmsis_nn_bias_data *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const cmsis_nn_bias_data *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 958,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L958"
              },
              "summary": "Basic s16 convolution function."
            },
            {
              "description": "Pointwise s16 convolution function: no stride, no padding, no dilation.\n\n1. Supported framework: TensorFlow Lite micro",
              "examples": [],
              "id": "arm_convolve_1x1_s16_ns_np_nd",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_s16_ns_np_nd",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const cmsis_nn_bias_data *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_s16_ns_np_nd(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const cmsis_nn_bias_data *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 999,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L999"
              },
              "summary": "Pointwise s16 convolution function: no stride, no padding, no dilation."
            },
            {
              "description": "arm_convolve_s16_fast_small_kernel function. The kernel size is <=8\n\n1. Supported framework: TensorFlow Lite micro",
              "examples": [],
              "id": "arm_convolve_s16_fast_small_kernel",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s16_fast_small_kernel",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_s16_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). conv_params->input_offset : Not used conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Struct with optional bias data pointer. Bias data type can be int64 or int32 depending flag in struct.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const cmsis_nn_bias_data *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` if successful or `ARM_CMSIS_NN_ARG_ERROR` if incorrect arguments or `ARM_CMSIS_NN_NO_IMPL_ERROR`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_s16_fast_small_kernel(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const cmsis_nn_bias_data *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 1040,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1040"
              },
              "summary": "armconvolves16fastsmallkernel function."
            },
            {
              "description": "Get the required buffer size for s16 convolution function.\n\nThe dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and non-Helium legs use different formulas.",
              "examples": [],
              "id": "arm_convolve_s16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_s16_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN] where HK and WK are the spatial filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_convolve_s16_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 1065,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1065"
              },
              "summary": "Get the required buffer size for s16 convolution function."
            },
            {
              "description": "Fast s4 version for 1x1 convolution (non-square shape).\n\n- Supported framework : TensorFlow Lite Micro\n- The following constrains on the arguments apply\n  \n  1. conv_params->padding.w = conv_params->padding.h = 0\n  2. conv_params->stride.w = conv_params->stride.h = 1",
              "examples": [],
              "id": "arm_convolve_1x1_s4_fast",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_s4_fast",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_1x1_s4_fast_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer ,if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8 packed with 2x int4",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_s4_fast(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1098,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1098"
              },
              "summary": "Fast s4 version for 1x1 convolution (non-square shape)."
            },
            {
              "description": "s4 version for 1x1 convolution with support for non-unity stride values\n\n- Supported framework : TensorFlow Lite Micro\n- The following constrains on the arguments apply\n  \n  1. conv_params->padding.w = conv_params->padding.h = 0",
              "examples": [],
              "id": "arm_convolve_1x1_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_s4",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. None is required by this function.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8 packed with 2x int4",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_s4(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1138,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1138"
              },
              "summary": "s4 version for 1x1 convolution with support for non-unity stride values"
            },
            {
              "description": "Fast s8 version for 1x1 convolution (non-square shape).\n\n- Supported framework : TensorFlow Lite Micro\n- The following constrains on the arguments apply\n  \n  1. conv_params->padding.w = conv_params->padding.h = 0\n  2. conv_params->stride.w = conv_params->stride.h = 1",
              "examples": [],
              "id": "arm_convolve_1x1_s8_fast",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_s8_fast",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_1x1_s8_fast_get_buffer_size will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. This function reads the buffer contents only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. On other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. Note also that on an Arm Compiler build (__ARMCC_VERSION >= 6010050) with ARM_MATH_DSP and without ARM_MATH_MVEI, supplying ctx->buf selects a buffered path that never reads weight_sum_ctx. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_s8_fast(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1204,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1204"
              },
              "summary": "Fast s8 version for 1x1 convolution (non-square shape)."
            },
            {
              "description": "Get the required buffer size for arm_convolve_1x1_s4_fast.",
              "examples": [],
              "id": "arm_convolve_1x1_s4_fast_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_s4_fast_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns the required buffer size in bytes, or -1 if input_dims->c is negative. No build needs this scratch buffer, so every valid shape returns 0."
                }
              ],
              "signature": "int32_t arm_convolve_1x1_s4_fast_get_buffer_size(const cmsis_nn_dims *input_dims)",
              "source": {
                "line": 1225,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1225"
              },
              "summary": "Get the required buffer size for armconvolve1x1s4fast."
            },
            {
              "description": "Get the required buffer size for arm_convolve_1x1_s8_fast.",
              "examples": [],
              "id": "arm_convolve_1x1_s8_fast_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_s8_fast_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns the required buffer size in bytes, or -1 if input_dims->c is negative. On builds that need this scratch buffer it also returns -1 if the required size would not fit in an int32_t; other builds need no buffer and return 0."
                }
              ],
              "signature": "int32_t arm_convolve_1x1_s8_fast_get_buffer_size(const cmsis_nn_dims *input_dims)",
              "source": {
                "line": 1236,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1236"
              },
              "summary": "Get the required buffer size for armconvolve1x1s8fast."
            },
            {
              "description": "s8 version for 1x1 convolution with support for non-unity stride values\n\n- Supported framework : TensorFlow Lite Micro\n- The following constrains on the arguments apply\n  \n  1. conv_params->padding.w = conv_params->padding.h = 0",
              "examples": [],
              "id": "arm_convolve_1x1_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_s8",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. None is required by this function.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. This function reads the buffer contents only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. On other builds the buffer contents are unread and a NULL buf is accepted, but the context struct itself is still dereferenced, so weight_sum_ctx must be non-NULL on every build. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, 1, 1, C_IN]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1285,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1285"
              },
              "summary": "s8 version for 1x1 convolution with support for non-unity stride values"
            },
            {
              "description": "1xn convolution\n\n- Supported framework : TensorFlow Lite Micro\n- The following constraints on the arguments apply\n  \n  1. input_dims->h, filter_dims->h and output_dims->h equal 1, and conv_params->padding.h is 0\n  2. conv_params->dilation.w is 1 and conv_params->stride.w is positive\n  3. conv_params->stride.w * input_dims->c is a multiple of 4\n  4. conv_params->padding.w, input_dims->w and output_dims->w are not negative, and filter_dims->w is at least 1\n- Any horizontal padding and output width are handled, including an odd total padding, a filter wider than the input and a VALID layer whose stride leaves trailing input unused. On MVE builds the output columns whose window starts before or ends past the input read a padded copy of the input columns they span, staged in ctx; the other columns read the input in place.",
              "examples": [],
              "id": "arm_convolve_1_x_n_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_s8",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_1_x_n_s8_get_buffer_size will return the buffer_size if required. buf must not be NULL. On builds with the MVE extension (ARM_MATH_MVEI) a non-zero ctx->size smaller than the staging the layer needs is rejected with ARM_CMSIS_NN_ARG_ERROR. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. The contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR; on other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1357,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1357"
              },
              "summary": "1xn convolution"
            },
            {
              "description": "Pre-computes per-output-channel weight sums for a standard convolution.\n\n- Supported framework : TensorFlow Lite Micro\n- The buffer pointed to by `vector_sum_buf` must be at least `output_dims->c × sizeof(int32_t)` bytes. `arm_convolve_s8_get_weights_sum_size()` returns that size on builds that use the sums, 0 elsewhere, and -1 for an output_dims->c that is negative or too large to size.\n- Layout: one int32 per output channel, indexed 0..`output_dims->c - 1`. Entry j holds `lhs_offset * sum(weights of output channel j) + bias_data[j]`, i.e. the bias and the input-offset contribution folded together. For grouped convolution the entries run over all output channels, with the groups laid out consecutively.\n- This is the buffer the `weight_sum_ctx` parameter of the s8 convolution kernels carries. Those kernels currently treat it as an input they only read, so it has to be filled before the call - see the individual functions for what each one currently does on MVE and non-MVE builds.\n- Reuse and invalidation: the contents depend only on `rhs`, `bias_data` and `lhs_offset`. They do not depend on the activations, so a buffer stays valid across calls and across batches for as long as those three are unchanged - for a static model the sums can be computed once at load time rather than per inference. Recompute whenever the weights, the bias or the input offset change (for example on requantization or a weight reload). The buffer is sized by one layer's `output_dims->c` and is specific to that layer's weights, so it cannot be shared between layers; give each layer its own.\n- Returns `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed.",
              "examples": [],
              "id": "arm_convolve_weight_sum",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_weight_sum",
              "params": [
                {
                  "description": "Pointer to the buffer that will hold the weight sums.",
                  "direction": "out",
                  "name": "vector_sum_buf",
                  "type": "int32_t *"
                },
                {
                  "description": "Pointer to the filter weights. Data type: int8",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, KH, KW, C_IN]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input-offset added to every input element before MAC. Range: [-127, 128]",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Optional bias pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_ARG_ERROR` on invalid arguments, `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed and the buffer is left untouched, or `ARM_CMSIS_NN_SUCCESS` on success. Portable code should not treat the `NO_IMPL_ERROR` case as a failure."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_weight_sum(\n    int32_t *vector_sum_buf,\n    const int8_t *rhs,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    const int32_t lhs_offset,\n    const int32_t *bias_data\n)",
              "source": {
                "line": 1409,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1409"
              },
              "summary": "Pre-computes per-output-channel weight sums for a standard convolution."
            },
            {
              "description": "Pre-computes per-channel weight sums for a depthwise convolution.\n\n- Supported framework : TensorFlow Lite Micro\n- Layout: one int32 per channel, sized by `arm_convolve_s8_get_weights_sum_size()`. Entry j holds `bias_data[j] + lhs_offset * sum(kernel values of channel j)`.\n- Reuse and invalidation follow the same rules as `arm_convolve_weight_sum()`: the contents depend only on `rhs`, `bias_data` and `lhs_offset`, so they may be computed once and reused until one of those changes, and they are specific to a single layer.\n- Returns `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed.\n- Not interchangeable with `arm_convolve_weight_sum()`: this function walks the channel-interleaved depthwise layout `[1, KH, KW, C_OUT]` with a stride of C_OUT, whereas `arm_convolve_weight_sum()` sums contiguous runs of `KH * KW * C_IN` weights. The two agree only by coincidence. Several in-tree tests do fill a depthwise weight_sum_ctx with `arm_convolve_weight_sum()` and are still correct, for one of three unrelated reasons: `arm_depthwise_conv_wrapper_s8()` does not consume the buffer on that route at all (ch_mult != 1, batches != 1, or a dilation the optimized route does not take - see that function); the wrapper converts the layer to a regular convolution, so conv-style sums are what is wanted; or C_OUT is 1, which collapses the stride-C_OUT walk to a contiguous one and makes the two helpers compute identical values. None of those generalise, so do not read them as licence to substitute one helper for the other. Use this function wherever the sums are actually read.",
              "examples": [],
              "id": "arm_depthwise_convolve_weight_sum",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_convolve_weight_sum",
              "params": [
                {
                  "description": "Buffer to hold the computed weight sums.",
                  "direction": "out",
                  "name": "vector_sum_buf",
                  "type": "int32_t *"
                },
                {
                  "description": "Currently unused: the implementation does not read or write it on any build, so NULL is accepted. Retained for signature compatibility; if a real buffer is passed, the caller is expected to clear it for security reasons.",
                  "direction": "inout",
                  "name": "scratch_buf",
                  "type": "int8_t *"
                },
                {
                  "description": "Depthwise convolution weights. Data type: int8",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Depthwise-convolution parameters (stride, dilation, pad, etc.)",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, KH, KW, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input-offset applied before MAC. Range: [-127, 128]",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Optional bias pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_ARG_ERROR` on invalid arguments, `ARM_CMSIS_NN_NO_IMPL_ERROR` on builds without the MVE extension, where the sums are currently not consumed and the buffer is left untouched, or `ARM_CMSIS_NN_SUCCESS` on success. Portable code should not treat the `NO_IMPL_ERROR` case as a failure."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_convolve_weight_sum(\n    int32_t *vector_sum_buf,\n    int8_t *scratch_buf,\n    const int8_t *rhs,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    const int32_t lhs_offset,\n    const int32_t *bias_data\n)",
              "source": {
                "line": 1458,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1458"
              },
              "summary": "Pre-computes per-channel weight sums for a depthwise convolution."
            },
            {
              "description": "Optimised convolution for 1x1 output images (shape of BX1x1xC_OUT) for 8x8 computations.\n\n- Supported framework : TensorFlow Lite Micro\n- Optimised for Bx1×1xC output CNN layers.\n- Constraints:\n  \n  1. `output_dims->h` and `output_dims->w` must equal 1\n  2. `output_dims->c` is expected to be a multiple of 4 for best performance",
              "examples": [],
              "id": "arm_convolve_1x1_out_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_out_s8",
              "params": [
                {
                  "description": "Function context that supplies a scratch buffer for activation rearrangement. A NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. The buffer must hold one 4-byte-aligned GEMM row, that is round_up_4(filter_dims->h * filter_dims->w * filter_dims->c) bytes, as returned by `arm_convolve_1x1_out_s8_get_buffer_size()`. The requirement does not scale with the group count: the kernel rewinds its im2col cursor to the start of the buffer after each group. Setting ctx->size lets this function reject an undersized buffer with ARM_CMSIS_NN_ARG_ERROR; leaving it at zero opts out of that check, which is what TFLite Micro and derivatives do today. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-output-channel weight sums, supplied by the caller. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data and conv_params->input_offset are unchanged - see `arm_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_convolve_weight_sum()`, passing conv_params->input_offset as lhs_offset and the same bias_data given here. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. The contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR; on other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: on MVE it yields wrong output while still returning ARM_CMSIS_NN_SUCCESS, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (stride, dilation, pad, offsets). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantisation multipliers and shifts.",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input data. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, KH, KW, C_IN]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to filter data. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, 1, 1, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output data. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_ARG_ERROR` on bad args, or `ARM_CMSIS_NN_SUCCESS` on success."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1x1_out_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1523,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1523"
              },
              "summary": "Optimised convolution for 1x1 output images (shape of BX1x1xCOUT) for 8x8 computations."
            },
            {
              "description": "Get the required scratch buffer size for `arm_convolve_1x1_out_s8()`.\n\n:::note\nThe figure is independent of the group count. `arm_convolve_1x1_out_s8()` rewinds its im2col cursor to the start of the buffer after each group's matmul, so groups do not accumulate.\n\n:::\n\n:::note\nCallers reaching the kernel through `arm_convolve_wrapper_s8()` must size the buffer with `arm_convolve_wrapper_s8_get_buffer_size()` instead, which covers every kernel the wrapper may dispatch to. This function is for callers that invoke `arm_convolve_1x1_out_s8()` directly.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1x1_out_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1x1_out_s8_get_buffer_size",
              "params": [
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, KH, KW, C_IN]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "For valid (non-negative, in-range) filter dimensions, the buffer size in bytes: round_up_4(KH * KW * C_IN) on builds with the MVE extension (ARM_MATH_MVEI), 0 otherwise, since `arm_convolve_1x1_out_s8()` only exists on MVE builds. Returns -1 if any of filter_dims->w, filter_dims->h or filter_dims->c is negative or out of int32_t range, or if the rounded-up product exceeds INT32_MAX. The validation runs on every build target, not just the MVE leg, so the contract does not vary by target."
                }
              ],
              "signature": "int32_t arm_convolve_1x1_out_s8_get_buffer_size(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 1554,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1554"
              },
              "summary": "Get the required scratch buffer size for armconvolve1x1outs8()."
            },
            {
              "description": "1xn convolution for s4 weights\n\n- Supported framework : TensorFlow Lite Micro\n- The following constrains on the arguments apply\n  \n  1. stride.w * input_dims->c is a multiple of 4\n  2. Explicit constraints(since it is for 1xN convolution) -## input_dims->h equals 1 -## output_dims->h equals 1 -## filter_dims->h equals 1\n\n:::note[Todo]\nRemove constraint on output_dims->w to make the function generic.\n\n:::",
              "examples": [],
              "id": "arm_convolve_1_x_n_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_s4",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. arm_convolve_1_x_n_s4_get_buffer_size will return the buffer_size if required The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8 as packed int4",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Optional bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_s4(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1592,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1592"
              },
              "summary": "1xn convolution for s4 weights"
            },
            {
              "description": "Get the required additional buffer size for 1xn convolution.",
              "examples": [],
              "id": "arm_convolve_1_x_n_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_s8_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or conv_params->stride.w is not positive. On builds with the MVE extension (ARM_MATH_MVEI) that is the staging size of `arm_convolve_1_x_n_s8()`, at least filter W * C_IN bytes, or -1 if it would not fit in an int32_t; other builds return `arm_convolve_s8_get_buffer_size()`."
                }
              ],
              "signature": "int32_t arm_convolve_1_x_n_s8_get_buffer_size(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1621,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1621"
              },
              "summary": "Get the required additional buffer size for 1xn convolution."
            },
            {
              "description": "Get the required additional buffer size for 1xn convolution.",
              "examples": [],
              "id": "arm_convolve_1_x_n_s4_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_convolve_1_x_n_s4_get_buffer_size",
              "params": [
                {
                  "description": "Convolution parameters (e.g. strides, dilations, pads,...). Range of conv_params->input_offset : [-127, 128] Range of conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, 1, WK, C_IN] where WK is the horizontal spatial filter dimension",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or conv_params->stride.w is not positive. It also returns -1 if the required size would not fit in an int32_t; on a Helium build the route whose padding lines up with the stride needs no buffer and returns 0 without computing one."
                }
              ],
              "signature": "int32_t arm_convolve_1_x_n_s4_get_buffer_size(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1643,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1643"
              },
              "summary": "Get the required additional buffer size for 1xn convolution."
            },
            {
              "description": "Wrapper function to pick the right optimized s8 depthwise convolution function.\n\n- Supported framework: TensorFlow Lite\n- Picks one of the the following functions\n  \n  1. `arm_depthwise_conv_s8()`\n  2. `arm_depthwise_conv_3x3_s8()` - Cortex-M CPUs with DSP extension only\n  3. `arm_depthwise_conv_s8_opt()`\n- Check details of `arm_depthwise_conv_s8_opt()` for potential data that can be accessed outside of the boundary.",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s8",
              "params": [
                {
                  "description": "Function context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s8_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s8_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-channel weight sums, supplied by the caller. The selected kernel only reads this buffer and never writes it, so it is filled once and may then be reused for as long as filter, bias and dw_conv_params->input_offset are unchanged - see `arm_depthwise_convolve_weight_sum()` for the layout and the full reuse rules. Whether the buffer is consumed at all depends on the route this wrapper takes. It is forwarded to `arm_depthwise_conv_s8_opt()`, which reads it under MVE, only when dw_conv_params->ch_mult == 1, input_dims->n == 1, and either both dilations are 1 or the layer is 1D and dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, dw_conv_params->padding.h == 0, dilation.h == 1 and dilation.w >= 1. Such a dilated 1D layer therefore reads the sums too. Outside those cases the wrapper calls `arm_depthwise_conv_s8()`, which has no such parameter and ignores the context entirely - which is why several in-tree tests legitimately pass sums built by `arm_convolve_weight_sum()`, or none at all, on those routes (a 2D-dilated layer, for example). On MVE with input_dims->c == 1 and an output channel count above CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (8 on armclang, 1 otherwise), the layer is instead converted to a regular convolution, and conv-style sums from `arm_convolve_weight_sum()` are what that route wants. Where the sums are actually read, fill the buffer with `arm_depthwise_convolve_weight_sum()`, passing dw_conv_params->input_offset as lhs_offset and the same bias given here, so that entry j holds input_offset * sum(weights of channel j) + bias[j]. That helper returns ARM_CMSIS_NN_NO_IMPL_ERROR on non-MVE builds, which is not a failure. Pass a valid context on every build. On the `arm_depthwise_conv_s8_opt()` route, a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on builds where the buffer is actually read (ARM_MATH_DSP and ARM_MATH_MVEI both defined); on other builds the parameter is unread and NULL is accepted. On MVE with input_dims->c == 1 and an output channel count above CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (8 on armclang, 1 otherwise), this wrapper instead diverts to `arm_convolve_wrapper_s8()`. That diversion exists only on MVE, and every kernel it can dispatch to diagnoses a NULL buf with ARM_CMSIS_NN_ARG_ERROR, so that route is covered too. None of this is a guarantee about future versions. Sized by `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions; see weight_sum_ctx for which dilated layers take the `arm_depthwise_conv_s8_opt()` route. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` on successful completion, or `ARM_CMSIS_NN_ARG_ERROR` on the `arm_depthwise_conv_s8_opt()` route if ctx->buf is NULL when a scratch buffer is required, or if weight_sum_ctx->buf is NULL on builds where it is read (ARM_MATH_DSP and ARM_MATH_MVEI both defined), or if ctx->size is non-zero and below `arm_depthwise_conv_s8_opt_get_buffer_size()` for a layer its channel path runs, or if that sizer returns -1 (a negative dimension or a byte count it cannot represent), or on the MVE `arm_convolve_wrapper_s8()` diversion route if weight_sum_ctx->buf is NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_wrapper_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1731,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1731"
              },
              "summary": "Wrapper function to pick the right optimized s8 depthwise convolution function."
            },
            {
              "description": "Wrapper function to pick the right optimized s4 depthwise convolution function.\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s4",
              "params": [
                {
                  "description": "Function context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s4_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s4_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer ,if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43].",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` - Successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_wrapper_s4(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1779,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1779"
              },
              "summary": "Wrapper function to pick the right optimized s4 depthwise convolution function."
            },
            {
              "description": "Get size of additional buffer required by `arm_depthwise_conv_wrapper_s8()`.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s8_get_buffer_size",
              "params": [
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1818,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1818"
              },
              "summary": "Get size of additional buffer required by armdepthwiseconvwrappers8()."
            },
            {
              "description": "Get size of additional buffer required by `arm_depthwise_conv_wrapper_s8()` for processors with DSP extension.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp",
              "params": [
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size_dsp(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1834,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1834"
              },
              "summary": "Get size of additional buffer required by armdepthwiseconvwrappers8() for processors with DSP extension."
            },
            {
              "description": "Get size of additional buffer required by `arm_depthwise_conv_wrapper_s8()` for Arm(R) Helium Architecture case.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. Which routes compute a byte count is build-dependent - a shape that does not select an optimized depthwise route, and the 3x3 route on builds without the MVE extension, need no buffer and short-circuit to 0 for any dimensions - so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s8_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s8_get_buffer_size_mve",
              "params": [
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. A route that needs no scratch buffer returns 0 for an in-range shape, but any route, including one that needs no buffer, may return -1 when a dimension it inspects is negative, so always test for -1 before using the value. A shape that does not select an optimized depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_s8_get_buffer_size_mve(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1850,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1850"
              },
              "summary": "Get size of additional buffer required by armdepthwiseconvwrappers8() for Arm(R) Helium Architecture case."
            },
            {
              "description": "Get size of additional buffer required by `arm_depthwise_conv_wrapper_s4()`.\n\nThis sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as `arm_depthwise_conv_s8_opt_get_buffer_size()`, so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s4_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s4_get_buffer_size",
              "params": [
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1878,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1878"
              },
              "summary": "Get size of additional buffer required by armdepthwiseconvwrappers4()."
            },
            {
              "description": "Get size of additional buffer required by `arm_depthwise_conv_wrapper_s4()` for processors with DSP extension.\n\nThis sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as `arm_depthwise_conv_s8_opt_get_buffer_size()`, so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s4_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked. This variant forwards to the top-level dispatcher, so it follows the build's leg; both legs inspect input_dims->c.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp",
              "params": [
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size_dsp(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1895,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1895"
              },
              "summary": "Get size of additional buffer required by armdepthwiseconvwrappers4() for processors with DSP extension."
            },
            {
              "description": "Get size of additional buffer required by `arm_depthwise_conv_wrapper_s4()` for Arm(R) Helium Architecture case.\n\nThis sizer routes straight to the s8 _mve/_dsp legs, both of which apply the same dimension check as `arm_depthwise_conv_s8_opt_get_buffer_size()`, so a negative input_dims->c, a negative filter dimension or an overflowing byte count is reported as -1 on every build target. A shape that does not select the optimized depthwise route short-circuits to 0 without the dimensions being range-checked, so a caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s4_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked. The Helium leg sizes its buffer from a fixed channel block rather than from input_dims->c, but it checks that dimension anyway so that this variant answers a negative channel count with the same -1 the dispatcher returns (issue #318).\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s4_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s4_get_buffer_size_mve",
              "params": [
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected leg reads is negative, or the required size would not fit in an int32_t. A shape that does not select the optimized depthwise route needs no buffer and returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_s4_get_buffer_size_mve(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1913,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1913"
              },
              "summary": "Get size of additional buffer required by armdepthwiseconvwrappers4() for Arm(R) Helium Architecture case."
            },
            {
              "description": "Basic s8 depthwise convolution function that doesn't have any constraints on the input dimensions.\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_depthwise_conv_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8",
              "params": [
                {
                  "description": "Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s8_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s8()`, size the context with `arm_depthwise_conv_wrapper_s8_get_buffer_size()` instead, because another route through that wrapper does require a buffer.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 1947,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1947"
              },
              "summary": "Basic s8 depthwise convolution function that doesn't have any constraints on the input dimensions."
            },
            {
              "description": "Basic s4 depthwise convolution function that doesn't have any constraints on the input dimensions.\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_depthwise_conv_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s4",
              "params": [
                {
                  "description": "Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s4_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s4()`, size the context with `arm_depthwise_conv_wrapper_s4_get_buffer_size()` instead, because another route through that wrapper does require a buffer.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43].",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "inout",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_s4(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output\n)",
              "source": {
                "line": 1989,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L1989"
              },
              "summary": "Basic s4 depthwise convolution function that doesn't have any constraints on the input dimensions."
            },
            {
              "description": "Basic s16 depthwise convolution function that doesn't have any constraints on the input dimensions.\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_depthwise_conv_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s16",
              "params": [
                {
                  "description": "Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_depthwise_conv_s16_get_buffer_size(). If you reached this kernel through `arm_depthwise_conv_wrapper_s16()`, size the context with `arm_depthwise_conv_wrapper_s16_get_buffer_size()` instead, because another route through that wrapper does require a buffer.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) conv_params->input_offset : Not used conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int64",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int64_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 2029,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2029"
              },
              "summary": "Basic s16 depthwise convolution function that doesn't have any constraints on the input dimensions."
            },
            {
              "description": "Wrapper function to pick the right optimized s16 depthwise convolution function.\n\n- Supported framework: TensorFlow Lite\n- Picks one of the the following functions\n  \n  1. `arm_depthwise_conv_s16()`\n  2. `arm_depthwise_conv_fast_s16()` - Cortex-M CPUs with DSP extension only",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s16",
              "params": [
                {
                  "description": "Function context (e.g. temporary buffer). Size ctx->buf with `arm_depthwise_conv_wrapper_s16_get_buffer_size()`, which accounts for whichever route this wrapper selects. It returns 0 when no buffer is needed, in which case ctx->buf may be NULL. `arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp()` and `arm_depthwise_conv_wrapper_s16_get_buffer_size_mve()` each size the buffer for one specific target and are not interchangeable. Use the variant matching the build the library is compiled for; when unsure, call the unsuffixed dispatching sizer, which always matches the wrapper's own routing. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Dilation is supported in both dimensions. When ch_mult == 1 and filter_dims->w * filter_dims->h < 512, `arm_depthwise_conv_fast_s16()` is used for an undilated layer and for a 1D layer dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, dw_conv_params->padding.h == 0, dilation.h == 1 and dilation.w >= 1. Other layers use `arm_depthwise_conv_s16()`. Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int64",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` - Successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_wrapper_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int64_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 2082,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2082"
              },
              "summary": "Wrapper function to pick the right optimized s16 depthwise convolution function."
            },
            {
              "description": "Get size of additional buffer required by `arm_depthwise_conv_wrapper_s16()`.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s16_get_buffer_size",
              "params": [
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2119,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2119"
              },
              "summary": "Get size of additional buffer required by armdepthwiseconvwrappers16()."
            },
            {
              "description": "Get size of additional buffer required by `arm_depthwise_conv_wrapper_s16()` for processors with DSP extension.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s16_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp",
              "params": [
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size_dsp(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2135,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2135"
              },
              "summary": "Get size of additional buffer required by armdepthwiseconvwrappers16() for processors with DSP extension."
            },
            {
              "description": "Get size of additional buffer required by `arm_depthwise_conv_wrapper_s16()` for Arm(R) Helium Architecture case.\n\nWhere a byte count is computed, an out-of-range shape is reported as -1 rather than a wrapped size, and the sentinel is propagated before any `ARM_NN_MAX()` or sum over sub-sizer results, so it can never collapse into a plausible positive size. A caller that needs its dimensions validated must validate them rather than infer validity from a non-negative return.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_wrapper_s16_get_buffer_size()`.\n\n:::\n\n:::note\nAn out-of-range shape is reported as -1, matching the top-level dispatcher, including the same caveat that a shape which needs no scratch buffer returns 0 without the dimensions being range-checked.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_wrapper_s16_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_wrapper_s16_get_buffer_size_mve",
              "params": [
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) Range of dw_conv_params->input_offset : Not used Range of dw_conv_params->input_offset : Not used",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN] Batch argument N is not used and assumed to be 1.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Size of additional memory required for optimizations in bytes, or -1 if the shape is out of range - a dimension the selected route reads is negative, or the required size would not fit in an int32_t. Only the fast depthwise route produces the -1, and it rejects a negative dimension on every build target, including the plain-C build where it needs no buffer - a route that needs no buffer may return -1, so always test for -1 before using the value. A shape that does not select the fast depthwise route returns 0 without range-checking the dimensions, so a 0 return is not a statement that the shape is valid."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_wrapper_s16_get_buffer_size_mve(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2152,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2152"
              },
              "summary": "Get size of additional buffer required by armdepthwiseconvwrappers16() for Arm(R) Helium Architecture case."
            },
            {
              "description": "Optimized s16 depthwise convolution function with constraint that in_channel equals out_channel.\n\n`ARM_CMSIS_NN_SUCCESS` - Successful operation\n\n- Supported framework: TensorFlow Lite\n- The following constraints on the arguments apply\n  \n  1. ch_mult == 1: the number of input channels equals the number of output channels\n  2. filter_dims->w * filter_dims->h < MAX_COL_COUNT (512)\n  3. dw_conv_params->dilation.h == 1 and dw_conv_params->dilation.w >= 1\n- Recommended when number of channels is 4 or greater.",
              "examples": [],
              "id": "arm_depthwise_conv_fast_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_fast_s16",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. `arm_depthwise_conv_fast_s16_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation.w is honoured; dw_conv_params->dilation.h must be 1. dw_conv_params->input_offset : Not used dw_conv_params->output_offset : Not used",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int64",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - ctx-buff == NULL and `arm_depthwise_conv_fast_s16_get_buffer_size()` != 0 or input channel != output channel or filter_dims->w * filter_dims->h >= MAX_COL_COUNT (512) or dw_conv_params->dilation.h != 1 or dw_conv_params->dilation.w < 1"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_fast_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int64_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 2200,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2200"
              },
              "summary": "Optimized s16 depthwise convolution function with constraint that inchannel equals outchannel."
            },
            {
              "description": "Get the required buffer size for optimized s16 depthwise convolution function with constraint that in_channel equals out_channel.\n\nThe dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.",
              "examples": [],
              "id": "arm_depthwise_conv_fast_s16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_fast_s16_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_depthwise_conv_fast_s16_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 2225,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2225"
              },
              "summary": "Get the required buffer size for optimized s16 depthwise convolution function with constraint that inchannel equals outchannel."
            },
            {
              "description": "Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below).\n\n- Supported framework : TensorFlow Lite Micro\n- The following constrains on the arguments apply\n  \n  1. Number of input channel equals number of output channels\n  2. Filter height and width equals 3\n  3. Padding along x is either 0 or 1.",
              "examples": [],
              "id": "arm_depthwise_conv_3x3_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_3x3_s8",
              "params": [
                {
                  "description": "Function context. This kernel uses no additional buffer, so ctx->buf may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - Unsupported dimension of tensors\n\n- Unsupported pad size along the x axis `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_3x3_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2261,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2261"
              },
              "summary": "Optimized s8 depthwise convolution function for 3x3 kernel size with some constraints on the input arguments(documented below)."
            },
            {
              "description": "Optimized s8 depthwise convolution function with constraint that in_channel equals out_channel.\n\n:::note\nThe second argument, weight_sum_ctx, has no counterpart on `arm_depthwise_conv_s8()`, so it is described here rather than by reference. It carries per-channel weight sums that the caller supplies: this function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as the weights, the bias and dw_conv_params->input_offset are unchanged - see `arm_depthwise_convolve_weight_sum()` for the layout and the full reuse rules. Fill it with `arm_depthwise_convolve_weight_sum()`, which walks the channel-interleaved depthwise weight layout; `arm_convolve_weight_sum()` sums a different set of weights and is not a substitute here. That helper returns `ARM_CMSIS_NN_NO_IMPL_ERROR` on non-MVE builds, which is not a failure. Size the buffer with `arm_convolve_s8_get_weights_sum_size()`: output_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise, and -1 for an output_dims->c that is negative or too large to size. Clear the buffer afterwards if applicable for security reasons. Pass a valid context on every build. On builds where the buffer is actually read (ARM_MATH_DSP and ARM_MATH_MVEI both defined), a NULL buf is diagnosed and this function returns `ARM_CMSIS_NN_ARG_ERROR`, matching `arm_convolve_s8()`. On other builds the parameter is unread and NULL is accepted. An allocated-but-unfilled buffer cannot be diagnosed the same way: it still produces wrong output while returning `ARM_CMSIS_NN_SUCCESS`, since an all-zero weight-sum vector is a legal result. None of this is a guarantee about future versions.\n\n:::\n\n:::note\nctx->size is optional: a caller that leaves it at zero opts out of the size check, as TFLM does. On the channel path, a non-zero ctx->size below `arm_depthwise_conv_s8_opt_get_buffer_size()` is rejected before any write. A layer the planar path takes needs only its plane, so it can succeed with less.\n\n:::\n\n:::note\nMVE channel tail loads and stores are predicated, so channel-indexed arrays are not accessed beyond the number of channels.\n\n:::\n\n- Supported framework: TensorFlow Lite\n- The following constrains on the arguments apply\n  \n  1. Number of input channel equals number of output channels or ch_mult equals 1\n- Reccomended when number of channels is 4 or greater.\n- On builds with ARM_MATH_DSP and ARM_MATH_MVEI, layers that `arm_depthwise_conv_s8_opt_planar_supported()` accepts run the planar path, with the same result as `arm_depthwise_conv_s8_opt_planar()`, unless ctx->size cannot hold its plane; every other layer runs the channel path of `arm_depthwise_conv_s8_opt_channelwise()`. Callers that choose the path ahead of time can call either one directly.",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt",
              "params": [
                {
                  "description": "Function context that contains the additional buffer if required by the function. `arm_depthwise_conv_s8_opt_get_buffer_size()` will return the buffer_size if required. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-channel weight sums, supplied by the caller and only read by this function. See the note below for how to size, fill and reuse the buffer and for when a NULL buf is diagnosed.",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation.w is honoured; dw_conv_params->dilation.h must be 1. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - input channel != output channel or ch_mult != 1, or dw_conv_params->dilation.h != 1 or dw_conv_params->dilation.w < 1, or ctx->buf is NULL when a scratch buffer is required, or ctx->size is non-zero and below `arm_depthwise_conv_s8_opt_get_buffer_size()` for a layer the channel path runs, or that sizer returns -1 (a negative dimension or a byte count it cannot represent) on the channel path, or weight_sum_ctx->buf is NULL on builds where it is read (ARM_MATH_DSP and ARM_MATH_MVEI both defined) `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_s8_opt(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2347,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2347"
              },
              "summary": "Optimized s8 depthwise convolution function with constraint that inchannel equals outchannel."
            },
            {
              "description": "Whether `arm_depthwise_conv_s8_opt()` runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels.\n\n- The rule is plain C and evaluates the same on every build, so a code generator can apply it ahead of time. The planar path itself exists only on builds with ARM_MATH_DSP and ARM_MATH_MVEI.\n- It depends only on the shapes and dw_conv_params: batch 1, C_IN equal to C_OUT, positive dimensions, stride 1, ch_mult 1, dilation.h 1, dilation.w at least 1 and at most 128 / C (integer division), at most 32 channels, the widths the path is faster for, and a plane that fits the scratch. This function is the reference for the rule; a mirror should be checked against it.\n- The width thresholds follow measured speed and may be retuned in a later release. A caller that calls `arm_depthwise_conv_s8_opt_planar()` directly must handle `ARM_CMSIS_NN_NO_IMPL_ERROR`, for example by calling `arm_depthwise_conv_s8_opt_channelwise()`.",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt_planar_supported",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt_planar_supported",
              "params": [
                {
                  "description": "Depthwise convolution parameters",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "1 when the planar path takes the layer with a scratch of `arm_depthwise_conv_s8_opt_get_buffer_size_mve()` bytes, 0 otherwise."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_s8_opt_planar_supported(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 2384,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2384"
              },
              "summary": "Whether armdepthwiseconvs8opt() runs a layer on its planar path, vectorized across the output pixels of each channel plane rather than across channels."
            },
            {
              "description": "The planar path of `arm_depthwise_conv_s8_opt()` on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane.\n\n- The output is identical to `arm_depthwise_conv_s8_opt()` and `arm_depthwise_conv_s8()`. The bias is read through the weight sums; bias_dims and bias_data are unused.",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt_planar",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt_planar",
              "params": [
                {
                  "description": "Function context with the `arm_depthwise_conv_s8_opt_get_buffer_size()` scratch",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()`",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()`",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_depthwise_conv_s8_opt()`, except its channel-path ctx->size check: a ctx->size too small for the plane returns ARM_CMSIS_NN_NO_IMPL_ERROR instead `ARM_CMSIS_NN_NO_IMPL_ERROR` - `arm_depthwise_conv_s8_opt_planar_supported()` rejects the layer, ctx->size cannot hold its plane, or the build lacks ARM_MATH_DSP or ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_s8_opt_planar(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2420,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2420"
              },
              "summary": "The planar path of armdepthwiseconvs8opt() on its own: s8 depthwise convolution vectorized across the output pixels of each channel plane."
            },
            {
              "description": "The channel-vectorized path of `arm_depthwise_conv_s8_opt()` on its own, without the planar attempt.\n\n- The output is identical to `arm_depthwise_conv_s8_opt()` and `arm_depthwise_conv_s8()` for every layer.",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt_channelwise",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt_channelwise",
              "params": [
                {
                  "description": "Function context with the `arm_depthwise_conv_s8_opt_get_buffer_size()` scratch",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()`",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()`",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - as for `arm_depthwise_conv_s8_opt()` `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_s8_opt_channelwise(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2457,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2457"
              },
              "summary": "The channel-vectorized path of armdepthwiseconvs8opt() on its own, without the planar attempt."
            },
            {
              "description": "s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of `arm_depthwise_conv_s8_opt()`, for layers in its gate. It computes three output rows per weight load.\n\n- The output is identical to `arm_depthwise_conv_s8_opt()` and `arm_depthwise_conv_s8()`. The bias is read through the weight sums, which `arm_depthwise_convolve_weight_sum()` fills as for `arm_depthwise_conv_s8_opt()`; bias_dims and bias_data are unused.\n- Gate: filter 3x3, dilation 1, N 1, C_IN equal to C_OUT with 16 <= C <= 2048 and C % 4 == 0, stride 1 or 2 and padding 0 or 1 in each dimension, input W >= 3 and H >= 1, output H >= 3 and W x H >= 16, every dimension at most 4096 and each tensor at most INT32_MAX elements, and the centre of the last output column's window inside the input: (output W - 1) x stride.w - padding.w + 1 < input W. Input rows above or below the input count as padding, as in `arm_depthwise_conv_s8()`.\n- Buffers: ctx and weight_sum_ctx and their buf are non-NULL, and ctx->size is at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` (3008 + input W x C + 16 bytes). ctx->buf needs no alignment.\n- It is a direct entry: `arm_depthwise_conv_s8_opt()` does not call it. A caller that selects the kernel per layer ahead of time calls `arm_depthwise_conv_s8_opt_3x3_c64_s1()` for C 64 with stride.h 1, this function for the rest of the gate, and `arm_depthwise_conv_s8_opt()` for every other layer, or on `ARM_CMSIS_NN_NO_IMPL_ERROR`. All three take the same arguments and the same weight sums; a ctx that serves all three holds the larger of `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` and `arm_depthwise_conv_s8_opt_get_buffer_size()`. On ARM_MATH_MVEI builds the second is the larger for a 3x3 filter when input W x C <= 1440.",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt_3x3",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt_3x3",
              "params": [
                {
                  "description": "Function context with at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` bytes of scratch",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()`",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()`",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_NO_IMPL_ERROR` - the layer or a buffer is outside the gate below, or the build lacks ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_s8_opt_3x3(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2513,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2513"
              },
              "summary": "s8 3x3 depthwise convolution that reads the input in place, without the im2col copy of armdepthwiseconvs8opt(), for layers in its gate."
            },
            {
              "description": "`arm_depthwise_conv_s8_opt_3x3()` specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps. It is the faster entry for those layers.\n\n- The output is identical to `arm_depthwise_conv_s8_opt_3x3()`, `arm_depthwise_conv_s8_opt()` and `arm_depthwise_conv_s8()`. The bias is read through the weight sums; bias_dims and bias_data are unused.\n- The two entries share only their gate, parameter packing and the code for the output_y % 3 remainder rows, so a build with -ffunction-sections and section garbage collection keeps only the code of the entries it calls.",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt_3x3_c64_s1",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt_3x3_c64_s1",
              "params": [
                {
                  "description": "Function context with at least `arm_depthwise_conv_s8_opt_3x3_get_buffer_size()` bytes of scratch",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-channel weight sums, as for `arm_depthwise_conv_s8_opt()`",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters, as for `arm_depthwise_conv_s8_opt()`",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_NO_IMPL_ERROR` - C_IN is not 64, dw_conv_params->stride.h is not 1, the layer or a buffer is outside the gate of `arm_depthwise_conv_s8_opt_3x3()`, or the build lacks ARM_MATH_MVEI; nothing is written `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_s8_opt_3x3_c64_s1(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2558,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2558"
              },
              "summary": "armdepthwiseconvs8opt3x3() specialized for 64 channels and a vertical stride of 1, with fixed channel offsets and edge columns that skip their zero taps."
            },
            {
              "description": "Get the scratch size in bytes of `arm_depthwise_conv_s8_opt_3x3()` and `arm_depthwise_conv_s8_opt_3x3_c64_s1()`.\n\n- The size is the minimum ctx->size both entries accept. It depends only on the input width and channel count, since the filter is always 3x3.\n- The function is plain C and returns the same size on every build, including builds without ARM_MATH_MVEI where the entries return `ARM_CMSIS_NN_NO_IMPL_ERROR`, so a code generator can size the scratch ahead of time. A non-negative size is not a statement that the layer is in the gate of the entries.",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt_3x3_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt_3x3_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]. Only W and C_IN are read.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "3008 + W x C_IN + 16 bytes, or -1 if W or C_IN is negative or the size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_depthwise_conv_s8_opt_3x3_get_buffer_size(const cmsis_nn_dims *input_dims)",
              "source": {
                "line": 2585,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2585"
              },
              "summary": "Get the scratch size in bytes of armdepthwiseconvs8opt3x3() and armdepthwiseconvs8opt3x3c64s1()."
            },
            {
              "description": "Optimized s4 depthwise convolution function with constraint that in_channel equals out_channel.\n\n:::note\nMVE channel tail loads and stores are predicated, so channel-indexed arrays are not accessed beyond the number of channels.\n\n:::\n\n- Supported framework: TensorFlow Lite\n- The following constrains on the arguments apply\n  \n  1. Number of input channel equals number of output channels or ch_mult equals 1\n- Reccomended when number of channels is 4 or greater.",
              "examples": [],
              "id": "arm_depthwise_conv_s4_opt",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s4_opt",
              "params": [
                {
                  "description": "Function context that contains the additional buffer required by the function. `arm_depthwise_conv_s4_opt_get_buffer_size()` will return the buffer_size. A NULL ctx->buf is diagnosed with `ARM_CMSIS_NN_ARG_ERROR`. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters (e.g. strides, dilations, pads,...) dw_conv_params->dilation is not used. Range of dw_conv_params->input_offset : [-127, 128] Range of dw_conv_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43].",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT]",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following `ARM_CMSIS_NN_ARG_ERROR` - input channel != output channel or ch_mult != 1 `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depthwise_conv_s4_opt(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2626,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2626"
              },
              "summary": "Optimized s4 depthwise convolution function with constraint that inchannel equals outchannel."
            },
            {
              "description": "Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel.\n\nThe dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_depthwise_conv_s8_opt_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 2651,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2651"
              },
              "summary": "Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel."
            },
            {
              "description": "Get the required buffer size for optimized s4 depthwise convolution function with constraint that in_channel equals out_channel.\n\nThe dimensions are not checked here: the query routes straight to the s8 _mve/_dsp leg and relies on the range checks inside that leg. Both legs apply the same check as `arm_depthwise_conv_s8_opt_get_buffer_size()`, so the answer for an out-of-range shape is the same on every build target.",
              "examples": [],
              "id": "arm_depthwise_conv_s4_opt_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s4_opt_get_buffer_size",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if input_dims->c or a filter dimension it reads is negative, or the required size would not fit in an int32_t."
                }
              ],
              "signature": "int32_t arm_depthwise_conv_s4_opt_get_buffer_size(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 2667,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2667"
              },
              "summary": "Get the required buffer size for optimized s4 depthwise convolution function with constraint that inchannel equals outchannel."
            },
            {
              "description": "Basic s4 Fully Connected function.\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_fully_connected_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_s4",
              "params": [
                {
                  "description": "Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_fully_connected_s4_get_buffer_size(). Do not size this context with `arm_fully_connected_s8_get_buffer_size()`: that sizes the kernel-sum buffer of a different kernel and does not describe this argument.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params *"
                },
                {
                  "description": "Per-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor.",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_tensor_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8_t packed 4-bit weights, e.g four sequential weights [0x1, 0x2, 0x3, 0x4] packed as [0x21, 0x43].",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_s4(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params *fc_params,\n    const cmsis_nn_per_tensor_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2717,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2717"
              },
              "summary": "Basic s4 Fully Connected function."
            },
            {
              "description": "Basic s8 Fully Connected function.\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_fully_connected_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_s8",
              "params": [
                {
                  "description": "Per-output-channel kernel sums, supplied by the caller - not scratch memory that this function fills in. The library never populates ctx->buf for this function, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds this function adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. Sized by `arm_fully_connected_s8_get_buffer_size()`: filter_dims->c * sizeof(int32_t) where the sums are used, 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params *"
                },
                {
                  "description": "Per-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor.",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_tensor_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params *fc_params,\n    const cmsis_nn_per_tensor_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2789,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2789"
              },
              "summary": "Basic s8 Fully Connected function."
            },
            {
              "description": "Basic s8 Fully Connected function using per channel quantization.\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_fully_connected_per_channel_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_per_channel_s8",
              "params": [
                {
                  "description": "Per-output-channel kernel sums, supplied by the caller - not scratch memory that this function fills in. The library never populates ctx->buf for this function, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. This function only reads the buffer and never writes it, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds this function adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. There is no per-channel sizer; `arm_fully_connected_s8_get_buffer_size()` returns the same quantity this function needs, filter_dims->c * sizeof(int32_t) where the sums are used and 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_per_channel_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params *fc_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2862,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2862"
              },
              "summary": "Basic s8 Fully Connected function using per channel quantization."
            },
            {
              "description": "s8 Fully Connected layer wrapper function\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_fully_connected_wrapper_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_wrapper_s8",
              "params": [
                {
                  "description": "Per-output-channel kernel sums, supplied by the caller - not scratch memory that this wrapper fills in. The library never populates ctx->buf here, and on builds with the MVE extension clearing it makes the output silently wrong, because the sums also carry the bias term. The context is passed straight through to `arm_fully_connected_per_channel_s8()` or `arm_fully_connected_s8()` depending on quant_params->is_per_channel, and both read it the same way. Fill it with `arm_vector_sum_s8()`, passing filter_dims->n as vector_cols, output_dims->c as vector_rows, filter_data as vector_data, fc_params->input_offset as lhs_offset, fc_params->filter_offset as rhs_offset and the same bias_data given here, so that entry j holds bias_data[j] plus input_offset times the sum of the weights of output channel j, where each of those filter_dims->n weights first has filter_offset added to it. That helper is available on every build and returns ARM_CMSIS_NN_SUCCESS. Neither selected kernel writes the buffer, so it is filled once and may then be reused for as long as filter_data, bias_data, fc_params->input_offset and fc_params->filter_offset are unchanged. Pass a valid context on every build; ctx itself is dereferenced unconditionally. Currently the contents are read only on builds with the MVE extension (ARM_MATH_MVEI), where the bias_data argument is ignored because the sums already carry it and a NULL buf is reported as ARM_CMSIS_NN_ARG_ERROR; an unfilled or cleared buffer there yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. On other builds the selected kernel adds bias_data directly and does not read ctx->buf. None of this is a guarantee about future versions. There is no wrapper sizer; `arm_fully_connected_s8_get_buffer_size()` returns the quantity both routes need, filter_dims->c * sizeof(int32_t) where the sums are used and 0 otherwise. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully Connected layer parameters. Range of fc_params->input_offset : [-127, 128] fc_params->filter_offset : 0 Range of fc_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params *"
                },
                {
                  "description": "Per-channel or per-tensor quantization info. Check struct defintion for details. It contains the multiplier and shift value(s) to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int32",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_wrapper_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params *fc_params,\n    const cmsis_nn_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 2937,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2937"
              },
              "summary": "s8 Fully Connected layer wrapper function"
            },
            {
              "description": "Calculate the sum of each row in vector_data, multiply by lhs_offset and optionally add s32 bias_data.",
              "examples": [],
              "id": "arm_vector_sum_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_vector_sum_s8",
              "params": [
                {
                  "description": "Buffer for vector sums",
                  "direction": "inout",
                  "name": "vector_sum_buf",
                  "type": "int32_t *"
                },
                {
                  "description": "Number of vector columns",
                  "direction": "in",
                  "name": "vector_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of vector rows",
                  "direction": "in",
                  "name": "vector_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Vector of weigths data",
                  "direction": "in",
                  "name": "vector_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Constant multiplied with each sum",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Constant added to each vector element before sum",
                  "direction": "in",
                  "name": "rhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Vector of bias data, added to each sum.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_vector_sum_s8(\n    int32_t *vector_sum_buf,\n    const int32_t vector_cols,\n    const int32_t vector_rows,\n    const int8_t *vector_data,\n    const int32_t lhs_offset,\n    const int32_t rhs_offset,\n    const int32_t *bias_data\n)",
              "source": {
                "line": 2961,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2961"
              },
              "summary": "Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s32 biasdata."
            },
            {
              "description": "Calculate the sum of each row in vector_data, multiply by lhs_offset and optionally add s64 bias_data.",
              "examples": [],
              "id": "arm_vector_sum_s8_s64",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_vector_sum_s8_s64",
              "params": [
                {
                  "description": "Buffer for vector sums",
                  "direction": "inout",
                  "name": "vector_sum_buf",
                  "type": "int64_t *"
                },
                {
                  "description": "Number of vector columns",
                  "direction": "in",
                  "name": "vector_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of vector rows",
                  "direction": "in",
                  "name": "vector_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Vector of weigths data",
                  "direction": "in",
                  "name": "vector_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Constant multiplied with each sum",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Vector of bias data, added to each sum.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int64_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_vector_sum_s8_s64(\n    int64_t *vector_sum_buf,\n    const int32_t vector_cols,\n    const int32_t vector_rows,\n    const int8_t *vector_data,\n    const int32_t lhs_offset,\n    const int64_t *bias_data\n)",
              "source": {
                "line": 2980,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2980"
              },
              "summary": "Calculate the sum of each row in vectordata, multiply by lhsoffset and optionally add s64 biasdata."
            },
            {
              "description": "Get size of additional buffer required by `arm_fully_connected_s8()`. See also arm_vector_sum_s8, which is required if buffer size is > 0.\n\nFor a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.",
              "examples": [],
              "id": "arm_fully_connected_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_s8_get_buffer_size",
              "params": [
                {
                  "description": "dimension of filter",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_fully_connected_s8_get_buffer_size(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 2998,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L2998"
              },
              "summary": "Get size of additional buffer required by armfullyconnecteds8()."
            },
            {
              "description": "Get size of additional buffer required by `arm_fully_connected_s8()` for processors with DSP extension.\n\nFor a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAlso validates dims like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_fully_connected_s8_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_s8_get_buffer_size_dsp",
              "params": [
                {
                  "description": "dimension of filter",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_fully_connected_s8_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 3009,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3009"
              },
              "summary": "Get size of additional buffer required by armfullyconnecteds8() for processors with DSP extension."
            },
            {
              "description": "Get size of additional buffer required by `arm_fully_connected_s8()` for Arm(R) Helium Architecture case.\n\nFor a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid filter_dims->c, returns -1 on every build target.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAlso validates dims like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_fully_connected_s8_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_s8_get_buffer_size_mve",
              "params": [
                {
                  "description": "dimension of filter",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_fully_connected_s8_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 3020,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3020"
              },
              "summary": "Get size of additional buffer required by armfullyconnecteds8() for Arm(R) Helium Architecture case."
            },
            {
              "description": "Basic s16 Fully Connected function.\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_fully_connected_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_s16",
              "params": [
                {
                  "description": "Unused. This function currently ignores the context entirely on every build - it neither reads nor writes ctx->buf - and `arm_fully_connected_s16_get_buffer_size()` returns 0 accordingly, so { NULL, 0 } is accepted. Unlike the s8 variants, no precomputed kernel sums are required here. None of this is a guarantee about future versions.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully Connected layer parameters. fc_params->input_offset : 0 fc_params->filter_offset : 0 fc_params->output_offset : 0",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params *"
                },
                {
                  "description": "Per-tensor quantization info. It contains the multiplier and shift value to be applied to the output tensor.",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_tensor_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int64",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params *fc_params,\n    const cmsis_nn_per_tensor_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int64_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 3057,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3057"
              },
              "summary": "Basic s16 Fully Connected function."
            },
            {
              "description": "Basic s16 Fully Connected function using per channel quantization.\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_fully_connected_per_channel_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_per_channel_s16",
              "params": [
                {
                  "description": "Scratch buffer that this function writes before it reads, on every build. It is filled here with one reduced int32 multiplier per output channel derived from quant_params->multiplier, so the caller supplies the storage only and the incoming contents are never used. Unlike the s8 variants, no precomputed kernel sums are expected, and clearing the buffer is harmless. Required on every build, not only under MVE: ctx or ctx->buf being NULL is reported as ARM_CMSIS_NN_ARG_ERROR, as is a non-zero ctx->size smaller than the requirement. A ctx->size of 0 is treated as undeclared and is not checked. Sized by `arm_fully_connected_per_channel_s16_get_buffer_size()`: filter_dims->c * sizeof(int32_t), which equals the output_dims->c entries written. The caller is expected to clear the buffer afterwards, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully Connected layer parameters. Range of fc_params->input_offset : 0 fc_params->filter_offset : 0 Range of fc_params->output_offset : 0",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params *"
                },
                {
                  "description": "Per-channel quantization info. It contains the multiplier and shift values to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int64",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_per_channel_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params *fc_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const int64_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 3114,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3114"
              },
              "summary": "Basic s16 Fully Connected function using per channel quantization."
            },
            {
              "description": "s16 Fully Connected layer wrapper function\n\n- Supported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_fully_connected_wrapper_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_wrapper_s16",
              "params": [
                {
                  "description": "Scratch buffer, whose use depends on the route taken. Unlike the s8 wrapper, no precomputed kernel sums are expected on either route, and clearing the buffer is harmless. When quant_params->is_per_channel is set, the context is passed to `arm_fully_connected_per_channel_s16()`, which writes it before reading it, on every build: it is filled there with one reduced int32 multiplier per output channel, so the caller supplies the storage only. On that route ctx or ctx->buf being NULL is reported as ARM_CMSIS_NN_ARG_ERROR, as is a non-zero ctx->size smaller than filter_dims->c * sizeof(int32_t); a ctx->size of 0 is treated as undeclared and is not checked. Size it with `arm_fully_connected_per_channel_s16_get_buffer_size()`. Otherwise the context goes to `arm_fully_connected_s16()`, which currently ignores it entirely, so { NULL, 0 } is accepted on that route. A caller that does not know the route in advance should size for the per-channel case, since `arm_fully_connected_s16_get_buffer_size()` returns 0. None of this is a guarantee about future versions. The caller is expected to clear the buffer afterwards, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Fully Connected layer parameters. Range of fc_params->input_offset : 0 fc_params->filter_offset : 0 Range of fc_params->output_offset : 0",
                  "direction": "in",
                  "name": "fc_params",
                  "type": "const cmsis_nn_fc_params *"
                },
                {
                  "description": "Per-channel or per-tensor quantization info. Check struct defintion for details. It contains the multiplier and shift value(s) to be applied to each output channel",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_quant_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN] Input dimension is taken as Nx(H * W * C_IN)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Two dimensional filter dimensions. Format: [N, C] N : accumulation depth and equals (H * W * C_IN) from input_dims C : output depth and equals C_OUT in output_dims H & W : Not used",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer. Data type: int8",
                  "direction": "in",
                  "name": "filter_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions. Format: [C_OUT] N, H, W : Not used",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Bias data pointer. Data type: int64",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, C_OUT] N : Batches C_OUT : Output depth H & W : Not used.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_fully_connected_wrapper_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_fc_params *fc_params,\n    const cmsis_nn_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const int64_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 3174,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3174"
              },
              "summary": "s16 Fully Connected layer wrapper function"
            },
            {
              "description": "Get size of additional buffer required by `arm_fully_connected_s16()`.",
              "examples": [],
              "id": "arm_fully_connected_s16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_s16_get_buffer_size",
              "params": [
                {
                  "description": "dimension of filter",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes"
                }
              ],
              "signature": "int32_t arm_fully_connected_s16_get_buffer_size(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 3192,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3192"
              },
              "summary": "Get size of additional buffer required by armfullyconnecteds16()."
            },
            {
              "description": "Get size of additional buffer required by `arm_fully_connected_s16()` for processors with DSP extension.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_s16_get_buffer_size()`.\n\n:::",
              "examples": [],
              "id": "arm_fully_connected_s16_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_s16_get_buffer_size_dsp",
              "params": [
                {
                  "description": "dimension of filter",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes"
                }
              ],
              "signature": "int32_t arm_fully_connected_s16_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 3202,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3202"
              },
              "summary": "Get size of additional buffer required by armfullyconnecteds16() for processors with DSP extension."
            },
            {
              "description": "Get size of additional buffer required by `arm_fully_connected_s16()` for Arm(R) Helium Architecture case.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_s16_get_buffer_size()`.\n\n:::",
              "examples": [],
              "id": "arm_fully_connected_s16_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_s16_get_buffer_size_mve",
              "params": [
                {
                  "description": "dimension of filter",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes"
                }
              ],
              "signature": "int32_t arm_fully_connected_s16_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 3212,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3212"
              },
              "summary": "Get size of additional buffer required by armfullyconnecteds16() for Arm(R) Helium Architecture case."
            },
            {
              "description": "Get size of additional buffer required by `arm_fully_connected_per_channel_s16()`.\n\nFor a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.",
              "examples": [],
              "id": "arm_fully_connected_per_channel_s16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_per_channel_s16_get_buffer_size",
              "params": [
                {
                  "description": "dimension of filter",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_fully_connected_per_channel_s16_get_buffer_size(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 3223,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3223"
              },
              "summary": "Get size of additional buffer required by armfullyconnectedperchannels16()."
            },
            {
              "description": "Get size of additional buffer required by `arm_fully_connected_per_channel_s16()` for processors with DSP extension.\n\nFor a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_per_channel_s16_get_buffer_size()`.\n\n:::\n\n:::note\nAlso validates dims like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_fully_connected_per_channel_s16_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_per_channel_s16_get_buffer_size_dsp",
              "params": [
                {
                  "description": "dimension of filter",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_fully_connected_per_channel_s16_get_buffer_size_dsp(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 3235,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3235"
              },
              "summary": "Get size of additional buffer required by armfullyconnectedperchannels16() for processors with DSP extension."
            },
            {
              "description": "Get size of additional buffer required by `arm_fully_connected_per_channel_s16()` for Arm(R) Helium Architecture case.\n\nFor a valid (non-negative, in-range) filter_dims->c, returns filter_dims->c * sizeof(int32_t) on every build target. For an invalid filter_dims->c, returns -1 on every build target.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_fully_connected_per_channel_s16_get_buffer_size()`.\n\n:::\n\n:::note\nAlso validates dims like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_fully_connected_per_channel_s16_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_fully_connected_per_channel_s16_get_buffer_size_mve",
              "params": [
                {
                  "description": "dimension of filter",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if filter_dims->c is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_fully_connected_per_channel_s16_get_buffer_size_mve(const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 3247,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3247"
              },
              "summary": "Get size of additional buffer required by armfullyconnectedperchannels16() for Arm(R) Helium Architecture case."
            },
            {
              "description": "s8 elementwise add of two tensors with support for broadcasting.",
              "examples": [],
              "id": "arm_add_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_add_s8",
              "params": [
                {
                  "description": "pointer to input tensor 1",
                  "direction": "in",
                  "name": "input1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input tensor 1 dimensions",
                  "direction": "in",
                  "name": "input1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to input tensor 2",
                  "direction": "in",
                  "name": "input2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input tensor 2 dimensions",
                  "direction": "in",
                  "name": "input2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "pointer to output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_add_s8(\n    const int8_t *input1_data,\n    const cmsis_nn_dims *input1_dims,\n    const int8_t *input2_data,\n    const cmsis_nn_dims *input2_dims,\n    const int32_t input1_offset,\n    const int32_t input1_mult,\n    const int32_t input1_shift,\n    const int32_t input2_offset,\n    const int32_t input2_mult,\n    const int32_t input2_shift,\n    const int32_t left_shift,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max\n)",
              "source": {
                "line": 3285,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3285"
              },
              "summary": "s8 elementwise add of two tensors with support for broadcasting."
            },
            {
              "description": "s8 elementwise add of scalar and vector",
              "examples": [],
              "id": "arm_add_scalar_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_add_scalar_s8",
              "params": [
                {
                  "description": "pointer to input scalar",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_add_scalar_s8(\n    const int8_t *input_1_vect,\n    const int8_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3329,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3329"
              },
              "summary": "s8 elementwise add of scalar and vector"
            },
            {
              "description": "s8 elementwise add of two vectors",
              "examples": [],
              "id": "arm_elementwise_add_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_add_s8",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "input left shift. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "inout",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_add_s8(\n    const int8_t *input_1_vect,\n    const int8_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3369,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3369"
              },
              "summary": "s8 elementwise add of two vectors"
            },
            {
              "description": "s8 elementwise absolute value",
              "examples": [],
              "id": "arm_abs_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_abs_s8",
              "params": [
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "input offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "indicates if output requantization is needed",
                  "direction": "in",
                  "name": "needs_rescale",
                  "type": "const bool"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_abs_s8(\n    const int8_t *input,\n    const int32_t input_offset,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const bool needs_rescale,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3400,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3400"
              },
              "summary": "s8 elementwise absolute value"
            },
            {
              "description": "s8 elementwise square root",
              "examples": [],
              "id": "arm_sqrt_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_sqrt_s8",
              "params": [
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "pointer to 256-entry lookup table",
                  "direction": "in",
                  "name": "sqrt_lut",
                  "type": "const int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_sqrt_s8(\n    const int8_t *input,\n    const cmsis_nn_dims *input_dims,\n    int8_t *output,\n    const int8_t *sqrt_lut\n)",
              "source": {
                "line": 3420,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3420"
              },
              "summary": "s8 elementwise square root"
            },
            {
              "description": "s16 elementwise square root using piecewise LUT with linear interpolation",
              "examples": [],
              "id": "arm_sqrt_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_sqrt_s16",
              "params": [
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "pointer to 513-entry lookup table (int16_t)",
                  "direction": "in",
                  "name": "sqrt_lut",
                  "type": "const int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_sqrt_s16(\n    const int16_t *input,\n    const cmsis_nn_dims *input_dims,\n    int16_t *output,\n    const int16_t *sqrt_lut\n)",
              "source": {
                "line": 3431,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3431"
              },
              "summary": "s16 elementwise square root using piecewise LUT with linear interpolation"
            },
            {
              "description": "s16 elementwise square root without a lookup table\n\nApproximates output[i] = trunc(sqrt(input[i] * scale)) saturated to 32767, which is LiteRT's int16 SQRT (dequantize in float32, sqrtf, divide by the output scale, truncate, clamp) for zero points 0, to within 1 LSB of LiteRT at every non-negative input for input scales 1e-7 to 1e-1 and output scales from 0.01x to 10x the full-range scale, saturating ones included. Inputs at or below 0 produce\n\n1. Needs no table; the int16 API does not depend on ARM_NN_ENABLE_F32/F16, and on targets without a floating-point unit the plain C path uses fmaf from the C library.",
              "examples": [],
              "id": "arm_sqrt_s16_tablefree",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_sqrt_s16_tablefree",
              "params": [
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "input_scale / (output_scale * output_scale) as float32: take the float32-rounded tensor scales, evaluate in float64 and round once to float32. Must be finite and greater than 0.",
                  "direction": "in",
                  "name": "scale",
                  "type": "const float"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_sqrt_s16_tablefree(\n    const int16_t *input,\n    const cmsis_nn_dims *input_dims,\n    int16_t *output,\n    const float scale\n)",
              "source": {
                "line": 3454,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3454"
              },
              "summary": "s16 elementwise square root without a lookup table"
            },
            {
              "description": "s16 elementwise absolute value",
              "examples": [],
              "id": "arm_abs_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_abs_s16",
              "params": [
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "input offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "indicates if output requantization is needed",
                  "direction": "in",
                  "name": "needs_rescale",
                  "type": "const bool"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_abs_s16(\n    const int16_t *input,\n    const int32_t input_offset,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const bool needs_rescale,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3470,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3470"
              },
              "summary": "s16 elementwise absolute value"
            },
            {
              "description": "INT16 reciprocal square root using a per-operator LUT.",
              "examples": [],
              "id": "arm_rsqrt_s16_per_op",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_rsqrt_s16_per_op",
              "params": [
                {
                  "description": "Pointer to the input buffer.",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor zero offset. The kernel evaluates each element as `input - input_offset` before the LUT lookup.",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer.",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor zero offset.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum output clamp.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum output clamp.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of elements.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to a 513-entry INT16 LUT in output domain.",
                  "direction": "in",
                  "name": "lut",
                  "type": "const int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_rsqrt_s16_per_op(\n    const int16_t *input,\n    const int32_t input_offset,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size,\n    const int16_t *lut\n)",
              "source": {
                "line": 3496,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3496"
              },
              "summary": "INT16 reciprocal square root using a per-operator LUT."
            },
            {
              "description": "INT16 reciprocal square root using a shared universal LUT.\n\nIn universal mode all RSQRT operators share a single LUT that captures the base 1/sqrt(x) shape, and operator-specific quantization is applied afterward via `out_mult` / `out_shift`. Because this two-step process introduces extra rounding stages, the output may differ from the per-op variant (`arm_rsqrt_s16_per_op`) by up to ±3 LSB per element. This is expected and acceptable for deployment.",
              "examples": [],
              "id": "arm_rsqrt_s16_universal",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_rsqrt_s16_universal",
              "params": [
                {
                  "description": "Pointer to the input buffer.",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor zero offset. The kernel evaluates each element as `input - input_offset` before the LUT lookup.",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer.",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor zero offset.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output requantization multiplier.",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Output requantization shift.",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Whether requantization is required.",
                  "direction": "in",
                  "name": "needs_rescale",
                  "type": "const bool"
                },
                {
                  "description": "Minimum output clamp.",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum output clamp.",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of elements.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to a 513-entry INT32 shared LUT in Q30 domain.",
                  "direction": "in",
                  "name": "lut",
                  "type": "const int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS or ARM_CMSIS_NN_ARG_ERROR."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_rsqrt_s16_universal(\n    const int16_t *input,\n    const int32_t input_offset,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const bool needs_rescale,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size,\n    const int32_t *lut\n)",
              "source": {
                "line": 3530,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3530"
              },
              "summary": "INT16 reciprocal square root using a shared universal LUT."
            },
            {
              "description": "s8 elementwise subtraction of two tensors with support for broadcasting.",
              "examples": [],
              "id": "arm_sub_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_sub_s8",
              "params": [
                {
                  "description": "pointer to input tensor 1",
                  "direction": "in",
                  "name": "input1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input tensor 1 dimensions",
                  "direction": "in",
                  "name": "input1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to input tensor 2",
                  "direction": "in",
                  "name": "input2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input tensor 2 dimensions",
                  "direction": "in",
                  "name": "input2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "pointer to output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_sub_s8(\n    const int8_t *input1_data,\n    const cmsis_nn_dims *input1_dims,\n    const int8_t *input2_data,\n    const cmsis_nn_dims *input2_dims,\n    const int32_t input1_offset,\n    const int32_t input1_mult,\n    const int32_t input1_shift,\n    const int32_t input2_offset,\n    const int32_t input2_mult,\n    const int32_t input2_shift,\n    const int32_t left_shift,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max\n)",
              "source": {
                "line": 3571,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3571"
              },
              "summary": "s8 elementwise subtraction of two tensors with support for broadcasting."
            },
            {
              "description": "s8 elementwise subtract of scalar and vector (scalar - vector)",
              "examples": [],
              "id": "arm_sub_scalar_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_sub_scalar_s8",
              "params": [
                {
                  "description": "pointer to input scalar",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "left shift applied to the result. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_sub_scalar_s8(\n    const int8_t *input_1_vect,\n    const int8_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3615,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3615"
              },
              "summary": "s8 elementwise subtract of scalar and vector (scalar - vector)"
            },
            {
              "description": "s8 elementwise subtract of two vectors",
              "examples": [],
              "id": "arm_elementwise_sub_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_sub_s8",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "input left shift. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "inout",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_sub_s8(\n    const int8_t *input_1_vect,\n    const int8_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3655,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3655"
              },
              "summary": "s8 elementwise subtract of two vectors"
            },
            {
              "description": "s16 elementwise add of two tensors with support for broadcasting.",
              "examples": [],
              "id": "arm_add_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_add_s16",
              "params": [
                {
                  "description": "pointer to input tensor 1",
                  "direction": "in",
                  "name": "input1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor 1 dimensions",
                  "direction": "in",
                  "name": "input1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to input tensor 2",
                  "direction": "in",
                  "name": "input2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor 2 dimensions",
                  "direction": "in",
                  "name": "input2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "pointer to output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_add_s16(\n    const int16_t *input1_data,\n    const cmsis_nn_dims *input1_dims,\n    const int16_t *input2_data,\n    const cmsis_nn_dims *input2_dims,\n    const int32_t input1_offset,\n    const int32_t input1_mult,\n    const int32_t input1_shift,\n    const int32_t input2_offset,\n    const int32_t input2_mult,\n    const int32_t input2_shift,\n    const int32_t left_shift,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max\n)",
              "source": {
                "line": 3702,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3702"
              },
              "summary": "s16 elementwise add of two tensors with support for broadcasting."
            },
            {
              "description": "s16 elementwise add of scalar and vector",
              "examples": [],
              "id": "arm_add_scalar_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_add_scalar_s16",
              "params": [
                {
                  "description": "pointer to input scalar",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "offset for input 1. Not used.",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Not used.",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset. Not used.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_add_scalar_s16(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3747,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3747"
              },
              "summary": "s16 elementwise add of scalar and vector"
            },
            {
              "description": "s16 elementwise add of two vectors",
              "examples": [],
              "id": "arm_elementwise_add_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_add_s16",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "offset for input 1. Not used.",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Not used.",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "input left shift. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "inout",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset. Not used.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_add_s16(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3789,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3789"
              },
              "summary": "s16 elementwise add of two vectors"
            },
            {
              "description": "s16 elementwise subtraction of two tensors with support for broadcasting.",
              "examples": [],
              "id": "arm_sub_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_sub_s16",
              "params": [
                {
                  "description": "pointer to input tensor 1",
                  "direction": "in",
                  "name": "input1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor 1 dimensions",
                  "direction": "in",
                  "name": "input1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to input tensor 2",
                  "direction": "in",
                  "name": "input2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor 2 dimensions",
                  "direction": "in",
                  "name": "input2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "pointer to output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_sub_s16(\n    const int16_t *input1_data,\n    const cmsis_nn_dims *input1_dims,\n    const int16_t *input2_data,\n    const cmsis_nn_dims *input2_dims,\n    const int32_t input1_offset,\n    const int32_t input1_mult,\n    const int32_t input1_shift,\n    const int32_t input2_offset,\n    const int32_t input2_mult,\n    const int32_t input2_shift,\n    const int32_t left_shift,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max\n)",
              "source": {
                "line": 3836,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3836"
              },
              "summary": "s16 elementwise subtraction of two tensors with support for broadcasting."
            },
            {
              "description": "s16 elementwise subtract of scalar and vector (scalar - vector)",
              "examples": [],
              "id": "arm_sub_scalar_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_sub_scalar_s16",
              "params": [
                {
                  "description": "pointer to input scalar",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "offset for input 1. Not used.",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Not used.",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "left shift applied to the result. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset. Not used.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_sub_scalar_s16(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3881,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3881"
              },
              "summary": "s16 elementwise subtract of scalar and vector (scalar - vector)"
            },
            {
              "description": "s16 elementwise subtract of two vectors",
              "examples": [],
              "id": "arm_elementwise_sub_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_sub_s16",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "offset for input 1. Not used.",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Not used.",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "input left shift. Bound: the kernel evaluates value << left_shift in int32; the offsets are unused, so with full-range int16 inputs the extremes are +32767 and -32768, and -32768 << 16 is exactly INT32_MIN, which makes 16 the last shift that stays representable. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "inout",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset. Not used.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_sub_s16(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3923,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3923"
              },
              "summary": "s16 elementwise subtract of two vectors"
            },
            {
              "description": "s8 elementwise squared difference of two tensors with support for broadcasting.",
              "examples": [],
              "id": "arm_squared_difference_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_squared_difference_s8",
              "params": [
                {
                  "description": "pointer to input tensor 1",
                  "direction": "in",
                  "name": "input1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input tensor 1 dimensions",
                  "direction": "in",
                  "name": "input1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to input tensor 2",
                  "direction": "in",
                  "name": "input2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input tensor 2 dimensions",
                  "direction": "in",
                  "name": "input2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "pointer to output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_squared_difference_s8(\n    const int8_t *input1_data,\n    const cmsis_nn_dims *input1_dims,\n    const int8_t *input2_data,\n    const cmsis_nn_dims *input2_dims,\n    const int32_t input1_offset,\n    const int32_t input1_mult,\n    const int32_t input1_shift,\n    const int32_t input2_offset,\n    const int32_t input2_mult,\n    const int32_t input2_shift,\n    const int32_t left_shift,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max\n)",
              "source": {
                "line": 3969,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L3969"
              },
              "summary": "s8 elementwise squared difference of two tensors with support for broadcasting."
            },
            {
              "description": "s8 elementwise squared difference of scalar and vector.",
              "examples": [],
              "id": "arm_squared_difference_scalar_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_squared_difference_scalar_s8",
              "params": [
                {
                  "description": "pointer to input scalar",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_squared_difference_scalar_s8(\n    const int8_t *input_1_vect,\n    const int8_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 4013,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4013"
              },
              "summary": "s8 elementwise squared difference of scalar and vector."
            },
            {
              "description": "s8 elementwise squared difference of two vectors.",
              "examples": [],
              "id": "arm_elementwise_squared_difference_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_squared_difference_s8",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_squared_difference_s8(\n    const int8_t *input_1_vect,\n    const int8_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 4055,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4055"
              },
              "summary": "s8 elementwise squared difference of two vectors."
            },
            {
              "description": "s16 elementwise squared difference of two tensors with support for broadcasting.",
              "examples": [],
              "id": "arm_squared_difference_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_squared_difference_s16",
              "params": [
                {
                  "description": "pointer to input tensor 1",
                  "direction": "in",
                  "name": "input1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor 1 dimensions",
                  "direction": "in",
                  "name": "input1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to input tensor 2",
                  "direction": "in",
                  "name": "input2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor 2 dimensions",
                  "direction": "in",
                  "name": "input2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "offset for input 1",
                  "direction": "in",
                  "name": "input1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2",
                  "direction": "in",
                  "name": "input2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "pointer to output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "output offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_squared_difference_s16(\n    const int16_t *input1_data,\n    const cmsis_nn_dims *input1_dims,\n    const int16_t *input2_data,\n    const cmsis_nn_dims *input2_dims,\n    const int32_t input1_offset,\n    const int32_t input1_mult,\n    const int32_t input1_shift,\n    const int32_t input2_offset,\n    const int32_t input2_mult,\n    const int32_t input2_shift,\n    const int32_t left_shift,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max\n)",
              "source": {
                "line": 4101,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4101"
              },
              "summary": "s16 elementwise squared difference of two tensors with support for broadcasting."
            },
            {
              "description": "s16 elementwise squared difference of scalar and vector.",
              "examples": [],
              "id": "arm_squared_difference_scalar_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_squared_difference_scalar_s16",
              "params": [
                {
                  "description": "pointer to input scalar",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "offset for input 1",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_squared_difference_scalar_s16(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 4145,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4145"
              },
              "summary": "s16 elementwise squared difference of scalar and vector."
            },
            {
              "description": "s16 elementwise squared difference of two vectors.",
              "examples": [],
              "id": "arm_elementwise_squared_difference_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_squared_difference_s16",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "offset for input 1",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 1",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 1",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "multiplier for input 2",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "shift for input 2",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift applied to both inputs before requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_squared_difference_s16(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 4187,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4187"
              },
              "summary": "s16 elementwise squared difference of two vectors."
            },
            {
              "description": "s8 elementwise multiplication of two tensors with support for broadcasting.",
              "examples": [],
              "id": "arm_mul_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_mul_s8",
              "params": [
                {
                  "description": "pointer to input tensor 1",
                  "direction": "in",
                  "name": "input1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input tensor 1 dimensions",
                  "direction": "in",
                  "name": "input1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to input tensor 2",
                  "direction": "in",
                  "name": "input2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input tensor 2 dimensions",
                  "direction": "in",
                  "name": "input2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "pointer to output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_mul_s8(\n    const int8_t *input1_data,\n    const cmsis_nn_dims *input1_dims,\n    const int8_t *input2_data,\n    const cmsis_nn_dims *input2_dims,\n    const int32_t input1_offset,\n    const int32_t input2_offset,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max\n)",
              "source": {
                "line": 4224,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4224"
              },
              "summary": "s8 elementwise multiplication of two tensors with support for broadcasting."
            },
            {
              "description": "s8 elementwise multiplication of scalar and vector",
              "examples": [],
              "id": "arm_mul_scalar_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_mul_scalar_s8",
              "params": [
                {
                  "description": "pointer to input scalar",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_mul_scalar_s8(\n    const int8_t *input_1_vect,\n    const int8_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_2_offset,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 4254,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4254"
              },
              "summary": "s8 elementwise multiplication of scalar and vector"
            },
            {
              "description": "s8 elementwise multiplication\n\nSupported framework: TensorFlow Lite micro",
              "examples": [],
              "id": "arm_elementwise_mul_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_mul_s8",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "offset for input 1. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Range: -127 to 128",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "inout",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset. Range: -128 to 127",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -128",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 127",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_mul_s8(\n    const int8_t *input_1_vect,\n    const int8_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_2_offset,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 4283,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4283"
              },
              "summary": "s8 elementwise multiplication"
            },
            {
              "description": "s16 elementwise multiplication of two tensors with support for broadcasting.",
              "examples": [],
              "id": "arm_mul_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_mul_s16",
              "params": [
                {
                  "description": "pointer to input tensor 1",
                  "direction": "in",
                  "name": "input1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor 1 dimensions",
                  "direction": "in",
                  "name": "input1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to input tensor 2",
                  "direction": "in",
                  "name": "input2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input tensor 2 dimensions",
                  "direction": "in",
                  "name": "input2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "offset for input 1. Not used.",
                  "direction": "in",
                  "name": "input1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Not used.",
                  "direction": "in",
                  "name": "input2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "pointer to output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "output offset. Not used.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_mul_s16(\n    const int16_t *input1_data,\n    const cmsis_nn_dims *input1_dims,\n    const int16_t *input2_data,\n    const cmsis_nn_dims *input2_dims,\n    const int32_t input1_offset,\n    const int32_t input2_offset,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max\n)",
              "source": {
                "line": 4315,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4315"
              },
              "summary": "s16 elementwise multiplication of two tensors with support for broadcasting."
            },
            {
              "description": "s16 elementwise multiplication of scalar and vector",
              "examples": [],
              "id": "arm_mul_scalar_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_mul_scalar_s16",
              "params": [
                {
                  "description": "pointer to input scalar",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "offset for input 1. Not used.",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Not used.",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset. Not used.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_mul_scalar_s16(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_2_offset,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 4345,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4345"
              },
              "summary": "s16 elementwise multiplication of scalar and vector"
            },
            {
              "description": "s16 elementwise multiplication\n\nSupported framework: TensorFlow Lite micro",
              "examples": [],
              "id": "arm_elementwise_mul_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_mul_s16",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "offset for input 1. Not used.",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Not used.",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "inout",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset. Not used.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_mul_s16(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_2_offset,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 4374,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4374"
              },
              "summary": "s16 elementwise multiplication"
            },
            {
              "description": "s8 elementwise minimum w/ support for broadcasting and scalar inputs.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_minimum_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_minimum_s8",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_minimum_s8(\n    const cmsis_nn_context *ctx,\n    const int8_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int8_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 4405,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4405"
              },
              "summary": "s8 elementwise minimum w/ support for broadcasting and scalar inputs."
            },
            {
              "description": "s8 elementwise maximum w/ support for broadcasting and scalar inputs.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_maximum_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_maximum_s8",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_maximum_s8(\n    const cmsis_nn_context *ctx,\n    const int8_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int8_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 4432,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4432"
              },
              "summary": "s8 elementwise maximum w/ support for broadcasting and scalar inputs."
            },
            {
              "description": "s16 elementwise minimum w/ support for broadcasting and scalar inputs.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_minimum_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_minimum_s16",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_minimum_s16(\n    const cmsis_nn_context *ctx,\n    const int16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 4459,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4459"
              },
              "summary": "s16 elementwise minimum w/ support for broadcasting and scalar inputs."
            },
            {
              "description": "s16 elementwise maximum w/ support for broadcasting and scalar inputs.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_maximum_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_maximum_s16",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_maximum_s16(\n    const cmsis_nn_context *ctx,\n    const int16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 4486,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4486"
              },
              "summary": "s16 elementwise maximum w/ support for broadcasting and scalar inputs."
            },
            {
              "description": "s8 elementwise comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_comparison_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_comparison_s8",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Comparison operation to perform",
                  "direction": "in",
                  "name": "operation",
                  "type": "arm_nn_compare_operation"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_comparison_s8(\n    const cmsis_nn_context *ctx,\n    const int8_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int8_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    arm_nn_compare_operation operation\n)",
              "source": {
                "line": 4527,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4527"
              },
              "summary": "s8 elementwise comparison with support for broadcasting."
            },
            {
              "description": "s16 elementwise comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_comparison_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_comparison_s16",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Comparison operation to perform",
                  "direction": "in",
                  "name": "operation",
                  "type": "arm_nn_compare_operation"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, the two input shapes are not broadcast-compatible, or the output shape is not their broadcast shape. `ctx` is unused and may be NULL."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_comparison_s16(\n    const cmsis_nn_context *ctx,\n    const int16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift,\n    arm_nn_compare_operation operation\n)",
              "source": {
                "line": 4570,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4570"
              },
              "summary": "s16 elementwise comparison with support for broadcasting."
            },
            {
              "description": "s8 elementwise equality comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_equal_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_equal_s8",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_equal_s8(\n    const cmsis_nn_context *ctx,\n    const int8_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int8_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4610,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4610"
              },
              "summary": "s8 elementwise equality comparison with support for broadcasting."
            },
            {
              "description": "s8 elementwise inequality comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_not_equal_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_not_equal_s8",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_not_equal_s8(\n    const cmsis_nn_context *ctx,\n    const int8_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int8_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4649,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4649"
              },
              "summary": "s8 elementwise inequality comparison with support for broadcasting."
            },
            {
              "description": "s8 elementwise greater-than comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_greater_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_greater_s8",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_greater_s8(\n    const cmsis_nn_context *ctx,\n    const int8_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int8_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4688,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4688"
              },
              "summary": "s8 elementwise greater-than comparison with support for broadcasting."
            },
            {
              "description": "s8 elementwise greater-or-equal comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_greater_equal_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_greater_equal_s8",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_greater_equal_s8(\n    const cmsis_nn_context *ctx,\n    const int8_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int8_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4727,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4727"
              },
              "summary": "s8 elementwise greater-or-equal comparison with support for broadcasting."
            },
            {
              "description": "s8 elementwise less-than comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_less_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_less_s8",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_less_s8(\n    const cmsis_nn_context *ctx,\n    const int8_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int8_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4766,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4766"
              },
              "summary": "s8 elementwise less-than comparison with support for broadcasting."
            },
            {
              "description": "s8 elementwise less-or-equal comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_less_equal_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_less_equal_s8",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int8 inputs and zero-points the widest operand is 255, so left_shift is at most 23. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s8()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_less_equal_s8(\n    const cmsis_nn_context *ctx,\n    const int8_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int8_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4805,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4805"
              },
              "summary": "s8 elementwise less-or-equal comparison with support for broadcasting."
            },
            {
              "description": "s16 elementwise equality comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_equal_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_equal_s16",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_equal_s16(\n    const cmsis_nn_context *ctx,\n    const int16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4844,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4844"
              },
              "summary": "s16 elementwise equality comparison with support for broadcasting."
            },
            {
              "description": "s16 elementwise inequality comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_not_equal_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_not_equal_s16",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_not_equal_s16(\n    const cmsis_nn_context *ctx,\n    const int16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4883,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4883"
              },
              "summary": "s16 elementwise inequality comparison with support for broadcasting."
            },
            {
              "description": "s16 elementwise greater-than comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_greater_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_greater_s16",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_greater_s16(\n    const cmsis_nn_context *ctx,\n    const int16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4922,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4922"
              },
              "summary": "s16 elementwise greater-than comparison with support for broadcasting."
            },
            {
              "description": "s16 elementwise greater-or-equal comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_greater_equal_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_greater_equal_s16",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_greater_equal_s16(\n    const cmsis_nn_context *ctx,\n    const int16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 4961,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L4961"
              },
              "summary": "s16 elementwise greater-or-equal comparison with support for broadcasting."
            },
            {
              "description": "s16 elementwise less-than comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_less_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_less_s16",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_less_s16(\n    const cmsis_nn_context *ctx,\n    const int16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 5000,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5000"
              },
              "summary": "s16 elementwise less-than comparison with support for broadcasting."
            },
            {
              "description": "s16 elementwise less-or-equal comparison with support for broadcasting.",
              "examples": [],
              "id": "arm_less_equal_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_less_equal_s16",
              "params": [
                {
                  "description": "Unused; may be NULL.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pointer to input1 tensor",
                  "direction": "in",
                  "name": "input_1_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input1 tensor dimensions",
                  "direction": "in",
                  "name": "input_1_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input2 tensor",
                  "direction": "in",
                  "name": "input_2_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input2 tensor dimensions",
                  "direction": "in",
                  "name": "input_2_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor (bool values)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "bool *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero-point for input1 tensor",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input1 tensor",
                  "direction": "in",
                  "name": "input_1_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input1 tensor",
                  "direction": "in",
                  "name": "input_1_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Zero-point for input2 tensor",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier for input2 tensor",
                  "direction": "in",
                  "name": "input_2_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for input2 tensor",
                  "direction": "in",
                  "name": "input_2_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Common left shift prior to requantization. Bound: the kernel evaluates (value + offset) << left_shift in int32; with full- range int16 inputs and a zero zero-point the widest operand is 32768, so left_shift is at most 16, and a non-zero zero-point lowers it. The scale 1 << left_shift is itself representable up to 30. Not validated by the kernel.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "As `arm_comparison_s16()`: ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR for invalid arguments."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_less_equal_s16(\n    const cmsis_nn_context *ctx,\n    const int16_t *input_1_data,\n    const cmsis_nn_dims *input_1_dims,\n    const int16_t *input_2_data,\n    const cmsis_nn_dims *input_2_dims,\n    bool *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t input_1_offset,\n    const int32_t input_1_mult,\n    const int32_t input_1_shift,\n    const int32_t input_2_offset,\n    const int32_t input_2_mult,\n    const int32_t input_2_shift,\n    const int32_t left_shift\n)",
              "source": {
                "line": 5039,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5039"
              },
              "summary": "s16 elementwise less-or-equal comparison with support for broadcasting."
            },
            {
              "description": "Q7 RELU function.",
              "examples": [],
              "id": "arm_relu_q7",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_relu_q7",
              "params": [
                {
                  "description": "pointer to input",
                  "direction": "inout",
                  "name": "data",
                  "type": "int8_t *"
                },
                {
                  "description": "number of elements",
                  "direction": "in",
                  "name": "size",
                  "type": "uint16_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_relu_q7(int8_t *data, uint16_t size)",
              "source": {
                "line": 5067,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5067"
              },
              "summary": "Q7 RELU function."
            },
            {
              "description": "Q7 RELU6 function.",
              "examples": [],
              "id": "arm_relu6_q7",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_relu6_q7",
              "params": [
                {
                  "description": "pointer to input",
                  "direction": "inout",
                  "name": "data",
                  "type": "int8_t *"
                },
                {
                  "description": "number of elements",
                  "direction": "in",
                  "name": "size",
                  "type": "uint16_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_relu6_q7(int8_t *data, uint16_t size)",
              "source": {
                "line": 5074,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5074"
              },
              "summary": "Q7 RELU6 function."
            },
            {
              "description": "Q15 RELU function.",
              "examples": [],
              "id": "arm_relu_q15",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_relu_q15",
              "params": [
                {
                  "description": "pointer to input",
                  "direction": "inout",
                  "name": "data",
                  "type": "int16_t *"
                },
                {
                  "description": "number of elements",
                  "direction": "in",
                  "name": "size",
                  "type": "uint16_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_relu_q15(int16_t *data, uint16_t size)",
              "source": {
                "line": 5081,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5081"
              },
              "summary": "Q15 RELU function."
            },
            {
              "description": "S8 clamp function.\n\nThis function clamps each element in the input tensor to the range. This can be useful for activations such as relu(0, 6), relu(-1,1), etc.",
              "examples": [],
              "id": "arm_clamp_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_clamp_s8",
              "params": [
                {
                  "description": "Pointer to input",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Minimum value to clamp to",
                  "direction": "in",
                  "name": "act_min",
                  "type": "const int8_t"
                },
                {
                  "description": "Maximum value to clamp to",
                  "direction": "in",
                  "name": "act_max",
                  "type": "const int8_t"
                },
                {
                  "description": "Pointer to output",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_clamp_s8(\n    const int8_t *input,\n    const int8_t act_min,\n    const int8_t act_max,\n    int8_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5095,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5095"
              },
              "summary": "S8 clamp function."
            },
            {
              "description": "S16 clamp function.\n\nThis function clamps each element in the input tensor to the range. This can be useful for activations such as relu(0, 6), relu(-1,1), etc.",
              "examples": [],
              "id": "arm_clamp_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_clamp_s16",
              "params": [
                {
                  "description": "Pointer to input",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Minimum value to clamp to",
                  "direction": "in",
                  "name": "act_min",
                  "type": "const int16_t"
                },
                {
                  "description": "Maximum value to clamp to",
                  "direction": "in",
                  "name": "act_max",
                  "type": "const int16_t"
                },
                {
                  "description": "Pointer to output",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_clamp_s16(\n    const int16_t *input,\n    const int16_t act_min,\n    const int16_t act_max,\n    int16_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5113,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5113"
              },
              "summary": "S16 clamp function."
            },
            {
              "description": "S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127.",
              "examples": [],
              "id": "arm_relu_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_relu_s8",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "output_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "output_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_relu_s8(\n    const int8_t *input,\n    const int32_t input_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier,\n    const int32_t output_shift,\n    int8_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5132,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5132"
              },
              "summary": "S8 ReLU activation function lower and upper bounds are quantized representations of 0 and 127."
            },
            {
              "description": "S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.",
              "examples": [],
              "id": "arm_relu_generic_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_relu_generic_s8",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "output_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "output_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to",
                  "direction": "in",
                  "name": "act_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to",
                  "direction": "in",
                  "name": "act_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_relu_generic_s8(\n    const int8_t *input,\n    const int32_t input_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier,\n    const int32_t output_shift,\n    const int32_t act_min,\n    const int32_t act_max,\n    int8_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5155,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5155"
              },
              "summary": "S8 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc."
            },
            {
              "description": "S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767.",
              "examples": [],
              "id": "arm_relu_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_relu_s16",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "output_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "output_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_relu_s16(\n    const int16_t *input,\n    const int32_t input_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier,\n    const int32_t output_shift,\n    int16_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5178,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5178"
              },
              "summary": "S16 ReLU activation function lower and upper bounds are quantized representations of 0 and 32767."
            },
            {
              "description": "S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc.",
              "examples": [],
              "id": "arm_relu_generic_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_relu_generic_s16",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "output_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "output_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to",
                  "direction": "in",
                  "name": "act_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to",
                  "direction": "in",
                  "name": "act_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_relu_generic_s16(\n    const int16_t *input,\n    const int32_t input_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier,\n    const int32_t output_shift,\n    const int32_t act_min,\n    const int32_t act_max,\n    int16_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5201,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5201"
              },
              "summary": "S16 ReLU activation function (generic version) This generic version allows to set custom lower and upper bounds to implement ReLU6, etc."
            },
            {
              "description": "S8 Leaky ReLU activation function.",
              "examples": [],
              "id": "arm_leaky_relu_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_leaky_relu_s8",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier for the alpha parameter",
                  "direction": "in",
                  "name": "output_multiplier_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift for the alpha parameter",
                  "direction": "in",
                  "name": "output_shift_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier for the identity parameter",
                  "direction": "in",
                  "name": "output_multiplier_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift for the identity parameter",
                  "direction": "in",
                  "name": "output_shift_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_leaky_relu_s8(\n    const int8_t *input,\n    const int32_t input_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier_alpha,\n    const int32_t output_shift_alpha,\n    const int32_t output_multiplier_identity,\n    const int32_t output_shift_identity,\n    int8_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5226,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5226"
              },
              "summary": "S8 Leaky ReLU activation function."
            },
            {
              "description": "S16 Leaky ReLU activation function.",
              "examples": [],
              "id": "arm_leaky_relu_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_leaky_relu_s16",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier for the alpha parameter",
                  "direction": "in",
                  "name": "output_multiplier_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift for the alpha parameter",
                  "direction": "in",
                  "name": "output_shift_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier for the identity parameter",
                  "direction": "in",
                  "name": "output_multiplier_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift for the identity parameter",
                  "direction": "in",
                  "name": "output_shift_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements in the input tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_leaky_relu_s16(\n    const int16_t *input,\n    const int32_t input_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier_alpha,\n    const int32_t output_shift_alpha,\n    const int32_t output_multiplier_identity,\n    const int32_t output_shift_identity,\n    int16_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5251,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5251"
              },
              "summary": "S16 Leaky ReLU activation function."
            },
            {
              "description": "Logistic activation function for s16.",
              "examples": [],
              "id": "arm_logistic_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_logistic_s16",
              "params": [
                {
                  "description": "Pointer to the input tensor",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements in the input tensor",
                  "direction": "in",
                  "name": "input_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization multiplier",
                  "direction": "in",
                  "name": "input_multiplier",
                  "type": "int32_t"
                },
                {
                  "description": "Input quantization shift within the range [0, 31]",
                  "direction": "in",
                  "name": "input_left_shift",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_logistic_s16(\n    const int16_t *input,\n    int16_t *output,\n    const int32_t input_size,\n    int32_t input_multiplier,\n    int32_t input_left_shift\n)",
              "source": {
                "line": 5272,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5272"
              },
              "summary": "Logistic activation function for s16."
            },
            {
              "description": "Tanh activation function for s16.",
              "examples": [],
              "id": "arm_tanh_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_tanh_s16",
              "params": [
                {
                  "description": "Pointer to the input tensor",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements in the input tensor",
                  "direction": "in",
                  "name": "input_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization multiplier",
                  "direction": "in",
                  "name": "input_multiplier",
                  "type": "int32_t"
                },
                {
                  "description": "Input quantization shift within the range [0, 31]",
                  "direction": "in",
                  "name": "input_left_shift",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_tanh_s16(\n    const int16_t *input,\n    int16_t *output,\n    const int32_t input_size,\n    int32_t input_multiplier,\n    int32_t input_left_shift\n)",
              "source": {
                "line": 5289,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5289"
              },
              "summary": "Tanh activation function for s16."
            },
            {
              "description": "s16 neural network activation function using direct table look-up\n\nSupported framework: TensorFlow Lite for Microcontrollers. This activation function must be bit precise congruent with the corresponding TFLM tanh and sigmoid activation functions",
              "examples": [],
              "id": "arm_nn_activation_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_activation_s16",
              "params": [
                {
                  "description": "pointer to input data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to output",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "number of elements",
                  "direction": "in",
                  "name": "size",
                  "type": "const int32_t"
                },
                {
                  "description": "bit-width of the integer part, assumed to be smaller than 3.",
                  "direction": "in",
                  "name": "left_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "type of activation functions",
                  "direction": "in",
                  "name": "type",
                  "type": "const arm_nn_activation_type"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_activation_s16(\n    const int16_t *input,\n    int16_t *output,\n    const int32_t size,\n    const int32_t left_shift,\n    const arm_nn_activation_type type\n)",
              "source": {
                "line": 5308,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5308"
              },
              "summary": "s16 neural network activation function using direct table look-up"
            },
            {
              "description": "S8 Hard-Swish activation function (compatibility version).\n\nThis version is compatible with TFLite implementation of Hard-Swish. hires_input_scale = (1.0 / 128.0) * float(input_scale) relu_scale = 3.0 / 32768.0 out_mul_real = hires_input_scale / float(output_scale) relu_mul_real = hires_input_scale / relu_scale output_multiplier_fp, output_multiplier_exp = to_q15_exp(out_mul_real) relu_multiplier_fp, relu_multiplier_exp = to_q15_exp(relu_mul_real) Here to_q15_exp quantizes to Q31 with a frexp exponent, then rounds and saturates the Q31 multiplier to Q15. For input_scale = output_scale = 0.125, the output pair is (16384, -6) and the ReLU pair is (21845, 4).",
              "examples": [],
              "id": "arm_hard_swish_compat_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_hard_swish_compat_s8",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier in fixed point format",
                  "direction": "in",
                  "name": "output_multiplier_fp",
                  "type": "const int32_t"
                },
                {
                  "description": "Exponent for output multiplier",
                  "direction": "in",
                  "name": "output_multiplier_exp",
                  "type": "const int32_t"
                },
                {
                  "description": "ReLU6 multiplier in fixed point format",
                  "direction": "in",
                  "name": "relu_multiplier_fp",
                  "type": "const int32_t"
                },
                {
                  "description": "Exponent for ReLU6 multiplier",
                  "direction": "in",
                  "name": "relu_multiplier_exp",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR if output_multiplier_exp is positive."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_hard_swish_compat_s8(\n    const int8_t *input,\n    const int32_t input_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier_fp,\n    const int32_t output_multiplier_exp,\n    const int32_t relu_multiplier_fp,\n    const int32_t relu_multiplier_exp,\n    int8_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5338,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5338"
              },
              "summary": "S8 Hard-Swish activation function (compatibility version)."
            },
            {
              "description": "S8 Hard-Swish activation function (precise version).\n\nThis version uses int32_t for intermediate computations to provide better accuracy. relu_q3, relu_q6 = round(3 / input_scale), round(6 / input_scale) prescale = min value to multiply input by to avoid overflow in intermediate computations M = ((input_scale**2) / (6.0 * output_scale)) * (1 << prescale) output_multiplier, output_shift = quantize_multiplier(M)",
              "examples": [],
              "id": "arm_hard_swish_precise_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_hard_swish_precise_s8",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "output_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "output_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "ReLU6 Q3 value",
                  "direction": "in",
                  "name": "relu_q3",
                  "type": "const int32_t"
                },
                {
                  "description": "ReLU6 Q6 value",
                  "direction": "in",
                  "name": "relu_q6",
                  "type": "const int32_t"
                },
                {
                  "description": "Prescale to apply to input",
                  "direction": "in",
                  "name": "prescale",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_hard_swish_precise_s8(\n    const int8_t *input,\n    const int32_t input_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier,\n    const int32_t output_shift,\n    const int32_t relu_q3,\n    const int32_t relu_q6,\n    const int32_t prescale,\n    int8_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5369,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5369"
              },
              "summary": "S8 Hard-Swish activation function (precise version)."
            },
            {
              "description": "S16 Hard-Swish activation function (precise version).\n\nThis version uses int32_t for intermediate computations to provide better accuracy. relu_q3, relu_q6 = round(3 / input_scale), round(6 / input_scale) prescale = min value to multiply input by to avoid overflow in intermediate computations M = ((input_scale**2) / (6.0 * output_scale)) * (1 << prescale) output_multiplier, output_shift = quantize_multiplier(M)",
              "examples": [],
              "id": "arm_hard_swish_precise_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_hard_swish_precise_s16",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "output_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "output_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "ReLU6 Q3 value",
                  "direction": "in",
                  "name": "relu_q3",
                  "type": "const int32_t"
                },
                {
                  "description": "ReLU6 Q6 value",
                  "direction": "in",
                  "name": "relu_q6",
                  "type": "const int32_t"
                },
                {
                  "description": "Prescale to apply to input",
                  "direction": "in",
                  "name": "prescale",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements in the tensor",
                  "direction": "in",
                  "name": "output_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_hard_swish_precise_s16(\n    const int16_t *input,\n    const int32_t input_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier,\n    const int32_t output_shift,\n    const int32_t relu_q3,\n    const int32_t relu_q6,\n    const int32_t prescale,\n    int16_t *output,\n    const int32_t output_size\n)",
              "source": {
                "line": 5401,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5401"
              },
              "summary": "S16 Hard-Swish activation function (precise version)."
            },
            {
              "description": "S8 PReLU activation function.",
              "examples": [],
              "id": "arm_prelu_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_prelu_s8",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Alpha tensor dimensions. Format: [N, H, W, C]",
                  "direction": "in",
                  "name": "alpha_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the alpha buffer",
                  "direction": "in",
                  "name": "alpha",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Alpha tensor zero offset",
                  "direction": "in",
                  "name": "alpha_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier 1",
                  "direction": "in",
                  "name": "output_multiplier_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift 1",
                  "direction": "in",
                  "name": "output_shift_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier 2",
                  "direction": "in",
                  "name": "output_multiplier_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift 2",
                  "direction": "in",
                  "name": "output_shift_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, alpha does not broadcast into the input, or the output dimensions do not match the input dimensions."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_prelu_s8(\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input,\n    const cmsis_nn_dims *alpha_dims,\n    const int8_t *alpha,\n    const int32_t input_offset,\n    const int32_t alpha_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier_identity,\n    const int32_t output_shift_identity,\n    const int32_t output_multiplier_alpha,\n    const int32_t output_shift_alpha,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output\n)",
              "source": {
                "line": 5432,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5432"
              },
              "summary": "S8 PReLU activation function."
            },
            {
              "description": "Elementwise S8 PReLU activation function.",
              "examples": [],
              "id": "arm_elementwise_prelu_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_prelu_s8",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the alpha buffer (same shape as input)",
                  "direction": "in",
                  "name": "alpha",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Alpha tensor zero offset",
                  "direction": "in",
                  "name": "alpha_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input >= 0",
                  "direction": "in",
                  "name": "output_multiplier_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input >= 0",
                  "direction": "in",
                  "name": "output_shift_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input < 0",
                  "direction": "in",
                  "name": "output_multiplier_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input < 0",
                  "direction": "in",
                  "name": "output_shift_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements to process",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_prelu_s8(\n    const int8_t *input,\n    const int8_t *alpha,\n    const int32_t input_offset,\n    const int32_t alpha_offset,\n    const int32_t out_offset,\n    const int32_t output_multiplier_identity,\n    const int32_t output_shift_identity,\n    const int32_t output_multiplier_alpha,\n    const int32_t output_shift_alpha,\n    int8_t *output,\n    const int32_t block_size\n)",
              "source": {
                "line": 5462,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5462"
              },
              "summary": "Elementwise S8 PReLU activation function."
            },
            {
              "description": "Scalar S8 PReLU activation function.",
              "examples": [],
              "id": "arm_prelu_scalar_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_prelu_scalar_s8",
              "params": [
                {
                  "description": "Pointer to the scalar buffer (single value)",
                  "direction": "in",
                  "name": "scalar_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the non-scalar buffer",
                  "direction": "in",
                  "name": "non_scalar_vect",
                  "type": "const int8_t *"
                },
                {
                  "description": "True if the scalar buffer holds the input value, false if it holds alpha",
                  "direction": "in",
                  "name": "scalar_is_input",
                  "type": "const bool"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Alpha tensor zero offset",
                  "direction": "in",
                  "name": "alpha_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input >= 0",
                  "direction": "in",
                  "name": "output_multiplier_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input >= 0",
                  "direction": "in",
                  "name": "output_shift_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input < 0",
                  "direction": "in",
                  "name": "output_multiplier_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input < 0",
                  "direction": "in",
                  "name": "output_shift_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements to process when the non-scalar vector is used",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_MATH_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_prelu_scalar_s8(\n    const int8_t *scalar_vect,\n    const int8_t *non_scalar_vect,\n    const bool scalar_is_input,\n    const int32_t input_offset,\n    const int32_t alpha_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier_identity,\n    const int32_t output_shift_identity,\n    const int32_t output_multiplier_alpha,\n    const int32_t output_shift_alpha,\n    int8_t *output,\n    const int32_t block_size\n)",
              "source": {
                "line": 5491,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5491"
              },
              "summary": "Scalar S8 PReLU activation function."
            },
            {
              "description": "S16 PReLU activation function.",
              "examples": [],
              "id": "arm_prelu_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_prelu_s16",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Alpha tensor dimensions. Format: [N, H, W, C]",
                  "direction": "in",
                  "name": "alpha_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the alpha buffer",
                  "direction": "in",
                  "name": "alpha",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Alpha tensor zero offset",
                  "direction": "in",
                  "name": "alpha_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input >= 0",
                  "direction": "in",
                  "name": "output_multiplier_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input >= 0",
                  "direction": "in",
                  "name": "output_shift_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input < 0",
                  "direction": "in",
                  "name": "output_multiplier_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input < 0",
                  "direction": "in",
                  "name": "output_shift_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR when a pointer is NULL, a dimension is not positive, alpha does not broadcast into the input, or the output dimensions do not match the input dimensions."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_prelu_s16(\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input,\n    const cmsis_nn_dims *alpha_dims,\n    const int16_t *alpha,\n    const int32_t input_offset,\n    const int32_t alpha_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier_identity,\n    const int32_t output_shift_identity,\n    const int32_t output_multiplier_alpha,\n    const int32_t output_shift_alpha,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output\n)",
              "source": {
                "line": 5524,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5524"
              },
              "summary": "S16 PReLU activation function."
            },
            {
              "description": "Elementwise S16 PReLU activation function.",
              "examples": [],
              "id": "arm_elementwise_prelu_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_prelu_s16",
              "params": [
                {
                  "description": "Pointer to the input buffer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to the alpha buffer (same shape as input)",
                  "direction": "in",
                  "name": "alpha",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Alpha tensor zero offset",
                  "direction": "in",
                  "name": "alpha_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input >= 0",
                  "direction": "in",
                  "name": "output_multiplier_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input >= 0",
                  "direction": "in",
                  "name": "output_shift_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input < 0",
                  "direction": "in",
                  "name": "output_multiplier_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input < 0",
                  "direction": "in",
                  "name": "output_shift_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements to process",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_prelu_s16(\n    const int16_t *input,\n    const int16_t *alpha,\n    const int32_t input_offset,\n    const int32_t alpha_offset,\n    const int32_t out_offset,\n    const int32_t output_multiplier_identity,\n    const int32_t output_shift_identity,\n    const int32_t output_multiplier_alpha,\n    const int32_t output_shift_alpha,\n    int16_t *output,\n    const int32_t block_size\n)",
              "source": {
                "line": 5554,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5554"
              },
              "summary": "Elementwise S16 PReLU activation function."
            },
            {
              "description": "Scalar S16 PReLU activation function.",
              "examples": [],
              "id": "arm_prelu_scalar_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_prelu_scalar_s16",
              "params": [
                {
                  "description": "Pointer to the scalar buffer (single value)",
                  "direction": "in",
                  "name": "scalar_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to the non-scalar buffer",
                  "direction": "in",
                  "name": "non_scalar_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "True if the scalar buffer holds the input value, false if it holds alpha",
                  "direction": "in",
                  "name": "scalar_is_input",
                  "type": "const bool"
                },
                {
                  "description": "Input tensor zero offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Alpha tensor zero offset",
                  "direction": "in",
                  "name": "alpha_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor zero offset",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input >= 0",
                  "direction": "in",
                  "name": "output_multiplier_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input >= 0",
                  "direction": "in",
                  "name": "output_shift_identity",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier when input < 0",
                  "direction": "in",
                  "name": "output_multiplier_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift when input < 0",
                  "direction": "in",
                  "name": "output_shift_alpha",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output buffer",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements to process when the non-scalar vector is used",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_prelu_scalar_s16(\n    const int16_t *scalar_vect,\n    const int16_t *non_scalar_vect,\n    const bool scalar_is_input,\n    const int32_t input_offset,\n    const int32_t alpha_offset,\n    const int32_t output_offset,\n    const int32_t output_multiplier_identity,\n    const int32_t output_shift_identity,\n    const int32_t output_multiplier_alpha,\n    const int32_t output_shift_alpha,\n    int16_t *output,\n    const int32_t block_size\n)",
              "source": {
                "line": 5583,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5583"
              },
              "summary": "Scalar S16 PReLU activation function."
            },
            {
              "description": "s8 average pooling function.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_avgpool_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avgpool_s8",
              "params": [
                {
                  "description": "Function context. Size ctx->buf with arm_avgpool_s8_get_buffer_size(output_dims->w, input_dims->c). Note that it takes the output width and the input channel count rather than a `cmsis_nn_dims`. `arm_avgpool_s8_get_buffer_size_dsp()` and `arm_avgpool_s8_get_buffer_size_mve()` size the same buffer for a specific target. It returns 0 where no buffer is needed, in which case ctx->buf may be NULL. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pooling parameters",
                  "direction": "in",
                  "name": "pool_params",
                  "type": "const cmsis_nn_pool_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [H, W] Argument N and C are not used.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns (an extent of 0 or less), which writes nothing and does not use ctx `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments, including a negative channel count, a pooling window that does not overlap the input, window positions (output index times stride minus padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t, or, on builds without MVE, a NULL ctx, or a NULL ctx->buf where the sizer asks for a buffer. Nothing is written to output_data then."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_avgpool_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_pool_params *pool_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 5638,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5638"
              },
              "summary": "s8 average pooling function."
            },
            {
              "description": "Get the required buffer size for S8 average pooling function.\n\nUnlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. `arm_avgpool_s8()` depends on that sentinel being non-zero, since it reads a non-zero size as \"ctx->buf is required\" before touching the accumulator buffer.",
              "examples": [],
              "id": "arm_avgpool_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avgpool_s8_get_buffer_size",
              "params": [
                {
                  "description": "output tensor dimension",
                  "direction": "in",
                  "name": "dim_dst_width",
                  "type": "const int"
                },
                {
                  "description": "number of input tensor channels",
                  "direction": "in",
                  "name": "ch_src",
                  "type": "const int"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_avgpool_s8_get_buffer_size(const int dim_dst_width, const int ch_src)",
              "source": {
                "line": 5659,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5659"
              },
              "summary": "Get the required buffer size for S8 average pooling function."
            },
            {
              "description": "Get the required buffer size for S8 average pooling function for processors with DSP extension.\n\nUnlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. `arm_avgpool_s8()` depends on that sentinel being non-zero, since it reads a non-zero size as \"ctx->buf is required\" before touching the accumulator buffer.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_avgpool_s8_get_buffer_size()`.\n\n:::\n\n:::note\nThis is the leg that computes a byte count, so it also validates ch_src like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_avgpool_s8_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avgpool_s8_get_buffer_size_dsp",
              "params": [
                {
                  "description": "output tensor dimension",
                  "direction": "in",
                  "name": "dim_dst_width",
                  "type": "const int"
                },
                {
                  "description": "number of input tensor channels",
                  "direction": "in",
                  "name": "ch_src",
                  "type": "const int"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_avgpool_s8_get_buffer_size_dsp(const int dim_dst_width, const int ch_src)",
              "source": {
                "line": 5671,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5671"
              },
              "summary": "Get the required buffer size for S8 average pooling function for processors with DSP extension."
            },
            {
              "description": "Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case.\n\nUnlike the fully connected and SVDF families, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target. `arm_avgpool_s8()` depends on that sentinel being non-zero, since it reads a non-zero size as \"ctx->buf is required\" before touching the accumulator buffer.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_avgpool_s8_get_buffer_size()`.\n\n:::\n\n:::note\nThis variant needs no buffer, so it returns 0 for every in-range shape. It still validates ch_src like the top-level dispatcher and the DSP leg, returning -1 for a negative ch_src or one whose byte count would not fit in an int32_t, so all three entry points answer an out-of-range shape alike.\n\n:::",
              "examples": [],
              "id": "arm_avgpool_s8_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avgpool_s8_get_buffer_size_mve",
              "params": [
                {
                  "description": "output tensor dimension",
                  "direction": "in",
                  "name": "dim_dst_width",
                  "type": "const int"
                },
                {
                  "description": "number of input tensor channels",
                  "direction": "in",
                  "name": "ch_src",
                  "type": "const int"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_avgpool_s8_get_buffer_size_mve(const int dim_dst_width, const int ch_src)",
              "source": {
                "line": 5684,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5684"
              },
              "summary": "Get the required buffer size for S8 average pooling function for Arm(R) Helium Architecture case."
            },
            {
              "description": "s16 average pooling function.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_avgpool_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avgpool_s16",
              "params": [
                {
                  "description": "Function context. Size ctx->buf with arm_avgpool_s16_get_buffer_size(output_dims->w, input_dims->c). Note that it takes the output width and the input channel count rather than a `cmsis_nn_dims`. `arm_avgpool_s16_get_buffer_size_dsp()` and `arm_avgpool_s16_get_buffer_size_mve()` size the same buffer for a specific target. It returns 0 where no buffer is needed, in which case ctx->buf may be NULL. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pooling parameters",
                  "direction": "in",
                  "name": "pool_params",
                  "type": "const cmsis_nn_pool_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [H, W] Argument N and C are not used.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns (an extent of 0 or less), which writes nothing and does not use ctx `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments, including a negative channel count, a pooling window that does not overlap the input, window positions (output index times stride minus padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t, or, on builds that use the buffer, a NULL ctx, or a NULL ctx->buf where the sizer asks for a buffer. Nothing is written to output_data then."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_avgpool_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_pool_params *pool_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output_data\n)",
              "source": {
                "line": 5721,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5721"
              },
              "summary": "s16 average pooling function."
            },
            {
              "description": "Get the required buffer size for S16 average pooling function.\n\nAs in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.",
              "examples": [],
              "id": "arm_avgpool_s16_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avgpool_s16_get_buffer_size",
              "params": [
                {
                  "description": "output tensor dimension",
                  "direction": "in",
                  "name": "dim_dst_width",
                  "type": "const int"
                },
                {
                  "description": "number of input tensor channels",
                  "direction": "in",
                  "name": "ch_src",
                  "type": "const int"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_avgpool_s16_get_buffer_size(const int dim_dst_width, const int ch_src)",
              "source": {
                "line": 5741,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5741"
              },
              "summary": "Get the required buffer size for S16 average pooling function."
            },
            {
              "description": "Get the required buffer size for S16 average pooling function for processors with DSP extension.\n\nAs in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_avgpool_s16_get_buffer_size()`.\n\n:::\n\n:::note\nThis is the leg that computes a byte count, so it also validates ch_src like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_avgpool_s16_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avgpool_s16_get_buffer_size_dsp",
              "params": [
                {
                  "description": "output tensor dimension",
                  "direction": "in",
                  "name": "dim_dst_width",
                  "type": "const int"
                },
                {
                  "description": "number of input tensor channels",
                  "direction": "in",
                  "name": "ch_src",
                  "type": "const int"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_avgpool_s16_get_buffer_size_dsp(const int dim_dst_width, const int ch_src)",
              "source": {
                "line": 5753,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5753"
              },
              "summary": "Get the required buffer size for S16 average pooling function for processors with DSP extension."
            },
            {
              "description": "Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case.\n\nAs in the s8 variant, it is the DSP leg that carries a byte count here: for a valid (non-negative, in-range) ch_src this returns ch_src * sizeof(int32_t) on builds with the DSP extension but no MVE, and 0 on builds with MVE and on plain-C builds. For an invalid ch_src it returns -1 on every build target.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_avgpool_s16_get_buffer_size()`.\n\n:::\n\n:::note\nThis variant needs no buffer, so it returns 0 for every in-range shape. It still validates ch_src like the top-level dispatcher and the DSP leg, returning -1 for a negative ch_src or one whose byte count would not fit in an int32_t, so all three entry points answer an out-of-range shape alike.\n\n:::",
              "examples": [],
              "id": "arm_avgpool_s16_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_avgpool_s16_get_buffer_size_mve",
              "params": [
                {
                  "description": "output tensor dimension",
                  "direction": "in",
                  "name": "dim_dst_width",
                  "type": "const int"
                },
                {
                  "description": "number of input tensor channels",
                  "direction": "in",
                  "name": "ch_src",
                  "type": "const int"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if ch_src is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_avgpool_s16_get_buffer_size_mve(const int dim_dst_width, const int ch_src)",
              "source": {
                "line": 5766,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5766"
              },
              "summary": "Get the required buffer size for S16 average pooling function for Arm(R) Helium Architecture case."
            },
            {
              "description": "s8 max pooling function.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_max_pool_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_max_pool_s8",
              "params": [
                {
                  "description": "Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_max_pool_s8_get_buffer_size(). Max pooling needs no accumulator scratch, unlike `arm_avgpool_s8()`, whose sizer does not describe this argument.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pooling parameters",
                  "direction": "in",
                  "name": "pool_params",
                  "type": "const cmsis_nn_pool_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. The input tensor must not overlap with the output tensor. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [H, W] Argument N and C are not used.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns, which writes nothing `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments: a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to output_data then."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_max_pool_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_pool_params *pool_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 5798,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5798"
              },
              "summary": "s8 max pooling function."
            },
            {
              "description": "s16 max pooling function.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_max_pool_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_max_pool_s16",
              "params": [
                {
                  "description": "Function context. This kernel uses no additional buffer, so ctx->buf may be NULL and there is deliberately no arm_max_pool_s16_get_buffer_size(). Max pooling needs no accumulator scratch, unlike `arm_avgpool_s16()`, whose sizer does not describe this argument.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Pooling parameters",
                  "direction": "in",
                  "name": "pool_params",
                  "type": "const cmsis_nn_pool_params *"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input (activation) data pointer. The input tensor must not overlap with the output tensor. Data type: int16",
                  "direction": "in",
                  "name": "src",
                  "type": "const int16_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [H, W] Argument N and C are not used.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [H, W, C_OUT] Argument N is not used. C_OUT equals C_IN.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "inout",
                  "name": "dst",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` - Successful operation, including an output with no rows or no columns, which writes nothing `ARM_CMSIS_NN_ARG_ERROR` - In case of invalid arguments: a batch count below 1, a pooling window that does not overlap the input, or window positions (output index * stride - padding, including one stride past the last window, plus the filter extent, and input size minus position) that do not fit in an int32_t. Nothing is written to dst then."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_max_pool_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_pool_params *pool_params,\n    const cmsis_nn_dims *input_dims,\n    const int16_t *src,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    int16_t *dst\n)",
              "source": {
                "line": 5836,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5836"
              },
              "summary": "s16 max pooling function."
            },
            {
              "description": "S8 softmax function.\n\n:::note\nSupported framework: TensorFlow Lite micro (bit-accurate)\n\n:::",
              "examples": [],
              "id": "arm_softmax_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_softmax_s8",
              "params": [
                {
                  "description": "Pointer to the input tensor",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Number of rows in the input tensor",
                  "direction": "in",
                  "name": "num_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of elements in each input row",
                  "direction": "in",
                  "name": "row_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization multiplier",
                  "direction": "in",
                  "name": "mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization shift within the range [0, 31]",
                  "direction": "in",
                  "name": "shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum difference with max in row. Used to check if the quantized exponential operation can be performed",
                  "direction": "in",
                  "name": "diff_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_softmax_s8(\n    const int8_t *input,\n    const int32_t num_rows,\n    const int32_t row_size,\n    const int32_t mult,\n    const int32_t shift,\n    const int32_t diff_min,\n    int8_t *output\n)",
              "source": {
                "line": 5864,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5864"
              },
              "summary": "S8 softmax function."
            },
            {
              "description": "S8 to s16 softmax function.\n\n:::note\nSupported framework: TensorFlow Lite micro (bit-accurate)\n\n:::",
              "examples": [],
              "id": "arm_softmax_s8_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_softmax_s8_s16",
              "params": [
                {
                  "description": "Pointer to the input tensor",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Number of rows in the input tensor",
                  "direction": "in",
                  "name": "num_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of elements in each input row",
                  "direction": "in",
                  "name": "row_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization multiplier",
                  "direction": "in",
                  "name": "mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization shift within the range [0, 31]",
                  "direction": "in",
                  "name": "shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum difference with max in row. Used to check if the quantized exponential operation can be performed",
                  "direction": "in",
                  "name": "diff_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_softmax_s8_s16(\n    const int8_t *input,\n    const int32_t num_rows,\n    const int32_t row_size,\n    const int32_t mult,\n    const int32_t shift,\n    const int32_t diff_min,\n    int16_t *output\n)",
              "source": {
                "line": 5886,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5886"
              },
              "summary": "S8 to s16 softmax function."
            },
            {
              "description": "S16 softmax function.\n\n:::note\nSupported framework: TensorFlow Lite micro (bit-accurate)\n\n:::",
              "examples": [],
              "id": "arm_softmax_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_softmax_s16",
              "params": [
                {
                  "description": "Pointer to the input tensor",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Number of rows in the input tensor",
                  "direction": "in",
                  "name": "num_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of elements in each input row",
                  "direction": "in",
                  "name": "row_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization multiplier",
                  "direction": "in",
                  "name": "mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization shift within the range [0, 31]",
                  "direction": "in",
                  "name": "shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Softmax s16 layer parameters with two pointers to LUTs speficied below. For indexing the high 9 bits are used and 7 remaining for interpolation. That means 512 entries for the 9-bit indexing and 1 extra for interpolation, i.e. 513 values for each LUT.\n\n- Lookup table for exp(x), where x uniform distributed between [-10.0 , 0.0]\n- Lookup table for 1 / (1 + x), where x uniform distributed between [0.0 , 1.0]",
                  "direction": "in",
                  "name": "softmax_params",
                  "type": "const cmsis_nn_softmax_lut_s16 *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_ARG_ERROR` Argument error check failed `ARM_CMSIS_NN_SUCCESS` - Successful operation"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_softmax_s16(\n    const int16_t *input,\n    const int32_t num_rows,\n    const int32_t row_size,\n    const int32_t mult,\n    const int32_t shift,\n    const cmsis_nn_softmax_lut_s16 *softmax_params,\n    int16_t *output\n)",
              "source": {
                "line": 5915,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5915"
              },
              "summary": "S16 softmax function."
            },
            {
              "description": "U8 softmax function.\n\n:::note\nSupported framework: TensorFlow Lite micro (bit-accurate)\n\n:::",
              "examples": [],
              "id": "arm_softmax_u8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_softmax_u8",
              "params": [
                {
                  "description": "Pointer to the input tensor",
                  "direction": "in",
                  "name": "input",
                  "type": "const uint8_t *"
                },
                {
                  "description": "Number of rows in the input tensor",
                  "direction": "in",
                  "name": "num_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of elements in each input row",
                  "direction": "in",
                  "name": "row_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization multiplier",
                  "direction": "in",
                  "name": "mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization shift within the range [0, 31]",
                  "direction": "in",
                  "name": "shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum difference with max in row. Used to check if the quantized exponential operation can be performed",
                  "direction": "in",
                  "name": "diff_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output",
                  "type": "uint8_t *"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_softmax_u8(\n    const uint8_t *input,\n    const int32_t num_rows,\n    const int32_t row_size,\n    const int32_t mult,\n    const int32_t shift,\n    const int32_t diff_min,\n    uint8_t *output\n)",
              "source": {
                "line": 5938,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5938"
              },
              "summary": "U8 softmax function."
            },
            {
              "description": "Reshape a s8 vector into another with different shape.\n\n:::note\nThe output is expected to be in a memory area that does not overlap with the input's\n\n:::",
              "examples": [],
              "id": "arm_reshape_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reshape_s8",
              "params": [
                {
                  "description": "points to the s8 input vector",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "points to the s8 output vector",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "total size of the input and output vectors in bytes",
                  "direction": "in",
                  "name": "total_size",
                  "type": "const uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_reshape_s8(const int8_t *input, int8_t *output, const uint32_t total_size)",
              "source": {
                "line": 5960,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5960"
              },
              "summary": "Reshape a s8 vector into another with different shape."
            },
            {
              "description": "Nearest neighbor resize function for s8 data.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_resize_nearest_neighbor_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_resize_nearest_neighbor_s8",
              "params": [
                {
                  "description": "Pointer to the context buffer. The buffer must hold at least (output_height + output_width) int32_t elements.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Resize parameters",
                  "direction": "in",
                  "name": "resize_params",
                  "type": "const cmsis_nn_resize_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C]",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Output size tensor dimensions",
                  "direction": "in",
                  "name": "output_size_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output size tensor data",
                  "direction": "in",
                  "name": "output_size_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C]",
                  "direction": "in",
                  "name": "output_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_resize_nearest_neighbor_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_resize_params *resize_params,\n    const cmsis_nn_dims *input_shape,\n    const int8_t *input_data,\n    const cmsis_nn_dims *output_size_shape,\n    const int32_t *output_size_data,\n    const cmsis_nn_dims *output_shape,\n    int8_t *output_data\n)",
              "source": {
                "line": 5982,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L5982"
              },
              "summary": "Nearest neighbor resize function for s8 data."
            },
            {
              "description": "Nearest neighbor resize function for s16 data.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_resize_nearest_neighbor_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_resize_nearest_neighbor_s16",
              "params": [
                {
                  "description": "Pointer to the context buffer. The buffer must hold at least (output_height + output_width) int32_t elements.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Resize parameters",
                  "direction": "in",
                  "name": "resize_params",
                  "type": "const cmsis_nn_resize_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C]",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Output size tensor dimensions",
                  "direction": "in",
                  "name": "output_size_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output size tensor data",
                  "direction": "in",
                  "name": "output_size_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C]",
                  "direction": "in",
                  "name": "output_shape",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_resize_nearest_neighbor_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_resize_params *resize_params,\n    const cmsis_nn_dims *input_shape,\n    const int16_t *input_data,\n    const cmsis_nn_dims *output_size_shape,\n    const int32_t *output_size_data,\n    const cmsis_nn_dims *output_shape,\n    int16_t *output_data\n)",
              "source": {
                "line": 6011,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6011"
              },
              "summary": "Nearest neighbor resize function for s16 data."
            },
            {
              "description": "Space to Depth function for s8 data type.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_space_to_depth_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_space_to_depth_s8",
              "params": [
                {
                  "description": "Pointer to the input tensor. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Block size for space to depth transformation",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output tensor. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H/block_size, W/block_size, C_IN*block_size*block_size]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_space_to_depth_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t block_size,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 6036,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6036"
              },
              "summary": "Space to Depth function for s8 data type."
            },
            {
              "description": "Space to Depth function for s16 data type.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_space_to_depth_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_space_to_depth_s16",
              "params": [
                {
                  "description": "Pointer to the input tensor. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Block size for space to depth transformation",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output tensor. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H/block_size, W/block_size, C_IN*block_size*block_size]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_space_to_depth_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t block_size,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 6058,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6058"
              },
              "summary": "Space to Depth function for s16 data type."
            },
            {
              "description": "Depth to Space function for s8 data type.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_depth_to_space_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depth_to_space_s8",
              "params": [
                {
                  "description": "Pointer to the input tensor. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Block size for depth to space transformation",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output tensor. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H*block_size, W*block_size, C_IN/(block_size*block_size)]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depth_to_space_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t block_size,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 6080,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6080"
              },
              "summary": "Depth to Space function for s8 data type."
            },
            {
              "description": "Depth to Space function for s16 data type.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_depth_to_space_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depth_to_space_s16",
              "params": [
                {
                  "description": "Pointer to the input tensor. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Block size for depth to space transformation",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to the output tensor. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H*block_size, W*block_size, C_IN/(block_size*block_size)]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_depth_to_space_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t block_size,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 6102,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6102"
              },
              "summary": "Depth to Space function for s16 data type."
            },
            {
              "description": "Space to Batch ND function for s8 data type.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_space_to_batch_nd_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_space_to_batch_nd_s8",
              "params": [
                {
                  "description": "Pointer to the input tensor. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Block shape for space to batch transformation",
                  "direction": "in",
                  "name": "block_shape",
                  "type": "const cmsis_nn_tile *"
                },
                {
                  "description": "Padding for height and width. Format: [n->top, h->left, w->bottom, c->right]",
                  "direction": "in",
                  "name": "pad",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N*block_shape[0]*block_shape[1], (H + pad_top + pad_bottom)/block_shape[0], (W + pad_left + pad_right)/block_shape[1], C_IN]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero offset for the output tensor",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_space_to_batch_nd_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_tile *block_shape,\n    const cmsis_nn_dims *pad,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t output_offset\n)",
              "source": {
                "line": 6126,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6126"
              },
              "summary": "Space to Batch ND function for s8 data type."
            },
            {
              "description": "Space to Batch ND function for s16 data type.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_space_to_batch_nd_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_space_to_batch_nd_s16",
              "params": [
                {
                  "description": "Pointer to the input tensor. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Block shape for space to batch transformation",
                  "direction": "in",
                  "name": "block_shape",
                  "type": "const cmsis_nn_tile *"
                },
                {
                  "description": "Padding for height and width. Format: [n->top, h->left, w->bottom, c->right]",
                  "direction": "in",
                  "name": "pad",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N*block_shape[0]*block_shape[1], (H + pad_top + pad_bottom)/block_shape[0], (W + pad_left + pad_right)/block_shape[1], C_IN]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Zero offset for the output tensor. NOT USED. Assume symmetric quantization for s16.",
                  "direction": "in",
                  "name": "output_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_space_to_batch_nd_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_tile *block_shape,\n    const cmsis_nn_dims *pad,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t output_offset\n)",
              "source": {
                "line": 6152,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6152"
              },
              "summary": "Space to Batch ND function for s16 data type."
            },
            {
              "description": "Batch to Space ND function for s8 data type.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_batch_to_space_nd_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_to_space_nd_s8",
              "params": [
                {
                  "description": "Pointer to the input tensor. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Block shape for batch to space transformation",
                  "direction": "in",
                  "name": "block_shape",
                  "type": "const cmsis_nn_tile *"
                },
                {
                  "description": "Cropping for height and width. Format: [n->top, h->left, w->bottom, c->right]",
                  "direction": "in",
                  "name": "crop",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N/(block_shape[0]*block_shape[1]), H*block_shape[0] - crop_top - crop_bottom, W*block_shape[1] - crop_left - crop_right, C_IN]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_batch_to_space_nd_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_tile *block_shape,\n    const cmsis_nn_dims *crop,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 6177,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6177"
              },
              "summary": "Batch to Space ND function for s8 data type."
            },
            {
              "description": "Batch to Space ND function for s16 data type.\n\n- Supported Framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_batch_to_space_nd_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_to_space_nd_s16",
              "params": [
                {
                  "description": "Pointer to the input tensor. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Block shape for batch to space transformation",
                  "direction": "in",
                  "name": "block_shape",
                  "type": "const cmsis_nn_tile *"
                },
                {
                  "description": "Cropping for height and width. Format: [n->top, h->left, w->bottom, c->right]",
                  "direction": "in",
                  "name": "crop",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N/(block_shape[0]*block_shape[1]), H*block_shape[0] - crop_top - crop_bottom, W*block_shape[1] - crop_left - crop_right, C_IN]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_batch_to_space_nd_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_tile *block_shape,\n    const cmsis_nn_dims *crop,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 6201,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6201"
              },
              "summary": "Batch to Space ND function for s16 data type."
            },
            {
              "description": "Basic transpose function.",
              "examples": [],
              "id": "arm_transpose_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_s8",
              "params": [
                {
                  "description": "Input (activation) data pointer. Data type: int8",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Output data pointer. Data type: int8",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *const"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Output tensor dimensions. Format may be arbitrary relative to input format. The output dimension will depend on the permutation dimensions. In other words the out dimensions are the result of applying the permutation to the input dimensions. The first transpose_params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[permutations[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Transpose parameters. Contains permutation dimensions. num_dims must be in [1, 4] and permutations must be a bijection over [0, num_dims - 1].",
                  "direction": "in",
                  "name": "transpose_params",
                  "type": "const cmsis_nn_transpose_params *const"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_s8(\n    const int8_t *input_data,\n    int8_t *const output_data,\n    const cmsis_nn_dims *const input_dims,\n    const cmsis_nn_dims *const output_dims,\n    const cmsis_nn_transpose_params *const transpose_params\n)",
              "source": {
                "line": 6235,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6235"
              },
              "summary": "Basic transpose function."
            },
            {
              "description": "Basic s16 transpose function.",
              "examples": [],
              "id": "arm_transpose_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_transpose_s16",
              "params": [
                {
                  "description": "Input (activation) data pointer. Data type: int16",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Output data pointer. Data type: int16",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *const"
                },
                {
                  "description": "Input (activation) tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Output tensor dimensions. Format may be arbitrary relative to input format. The output dimension will depend on the permutation dimensions. In other words the out dimensions are the result of applying the permutation to the input dimensions. The first transpose_params->num_dims fields, taken in the order [N, H, W, C], must satisfy output[i] == input[permutations[i]]; the function returns `ARM_CMSIS_NN_ARG_ERROR` and writes nothing if they do not.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Transpose parameters. Contains permutation dimensions. num_dims must be in [1, 4] and permutations must be a bijection over [0, num_dims - 1].",
                  "direction": "in",
                  "name": "transpose_params",
                  "type": "const cmsis_nn_transpose_params *const"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_transpose_s16(\n    const int16_t *input_data,\n    int16_t *const output_data,\n    const cmsis_nn_dims *const input_dims,\n    const cmsis_nn_dims *const output_dims,\n    const cmsis_nn_transpose_params *const transpose_params\n)",
              "source": {
                "line": 6263,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6263"
              },
              "summary": "Basic s16 transpose function."
            },
            {
              "description": "int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate. The argument offset_x will be used to store the input tensor in the correct position in the output tensor\n\ni.e. offset_x = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_x(&input[i], ..., &output, ..., ..., offset_x) offset_x += input_x[i] }\n\nThis function assumes that the output tensor has:\n\n1. The same height of the input tensor\n2. The same number of channels of the input tensor\n3. The same batch size of the input tensor\n\nUnless specified otherwise, arguments are mandatory.\n\n:::note\nThis function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation\n\n:::\n\n**Input constraints** offset_x is less than output_x",
              "examples": [],
              "id": "arm_concatenation_s8_x",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_s8_x",
              "params": [
                {
                  "description": "Pointer to input tensor. Input tensor must not overlap with the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Width of input tensor",
                  "direction": "in",
                  "name": "input_x",
                  "type": "const uint16_t"
                },
                {
                  "description": "Height of input tensor",
                  "direction": "in",
                  "name": "input_y",
                  "type": "const uint16_t"
                },
                {
                  "description": "Channels in input tensor",
                  "direction": "in",
                  "name": "input_z",
                  "type": "const uint16_t"
                },
                {
                  "description": "Batch size in input tensor",
                  "direction": "in",
                  "name": "input_w",
                  "type": "const uint16_t"
                },
                {
                  "description": "Pointer to output tensor. Expected to be at least (input_x * input_y * input_z * input_w) + offset_x bytes.",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Width of output tensor",
                  "direction": "in",
                  "name": "output_x",
                  "type": "const uint16_t"
                },
                {
                  "description": "The offset (in number of elements) on the X axis to start concatenating the input tensor It is user responsibility to provide the correct value",
                  "direction": "in",
                  "name": "offset_x",
                  "type": "const uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_s8_x(\n    const int8_t *input,\n    const uint16_t input_x,\n    const uint16_t input_y,\n    const uint16_t input_z,\n    const uint16_t input_w,\n    int8_t *output,\n    const uint16_t output_x,\n    const uint32_t offset_x\n)",
              "source": {
                "line": 6312,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6312"
              },
              "summary": "int8/uint8 concatenation function to be used for concatenating N-tensors along the X axis This function should be called for each input tensor to concatenate."
            },
            {
              "description": "int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate. The argument offset_y will be used to store the input tensor in the correct position in the output tensor\n\ni.e. offset_y = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_y(&input[i], ..., &output, ..., ..., offset_y) offset_y += input_y[i] }\n\nThis function assumes that the output tensor has:\n\n1. The same width of the input tensor\n2. The same number of channels of the input tensor\n3. The same batch size of the input tensor\n\nUnless specified otherwise, arguments are mandatory.\n\n:::note\nThis function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation\n\n:::\n\n**Input constraints** offset_y is less than output_y",
              "examples": [],
              "id": "arm_concatenation_s8_y",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_s8_y",
              "params": [
                {
                  "description": "Pointer to input tensor. Input tensor must not overlap with the output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Width of input tensor",
                  "direction": "in",
                  "name": "input_x",
                  "type": "const uint16_t"
                },
                {
                  "description": "Height of input tensor",
                  "direction": "in",
                  "name": "input_y",
                  "type": "const uint16_t"
                },
                {
                  "description": "Channels in input tensor",
                  "direction": "in",
                  "name": "input_z",
                  "type": "const uint16_t"
                },
                {
                  "description": "Batch size in input tensor",
                  "direction": "in",
                  "name": "input_w",
                  "type": "const uint16_t"
                },
                {
                  "description": "Pointer to output tensor. Expected to be at least (input_z * input_w * input_x * input_y) + offset_y bytes.",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Height of output tensor",
                  "direction": "in",
                  "name": "output_y",
                  "type": "const uint16_t"
                },
                {
                  "description": "The offset on the Y axis to start concatenating the input tensor It is user responsibility to provide the correct value",
                  "direction": "in",
                  "name": "offset_y",
                  "type": "const uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_s8_y(\n    const int8_t *input,\n    const uint16_t input_x,\n    const uint16_t input_y,\n    const uint16_t input_z,\n    const uint16_t input_w,\n    int8_t *output,\n    const uint16_t output_y,\n    const uint32_t offset_y\n)",
              "source": {
                "line": 6359,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6359"
              },
              "summary": "int8/uint8 concatenation function to be used for concatenating N-tensors along the Y axis This function should be called for each input tensor to concatenate."
            },
            {
              "description": "int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate. The argument offset_z will be used to store the input tensor in the correct position in the output tensor\n\ni.e. offset_z = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_z(&input[i], ..., &output, ..., ..., offset_z) offset_z += input_z[i] }\n\nThis function assumes that the output tensor has:\n\n1. The same width of the input tensor\n2. The same height of the input tensor\n3. The same batch size of the input tensor\n\nUnless specified otherwise, arguments are mandatory.\n\n:::note\nThis function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation\n\n:::\n\n**Input constraints** offset_z is less than output_z",
              "examples": [],
              "id": "arm_concatenation_s8_z",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_s8_z",
              "params": [
                {
                  "description": "Pointer to input tensor. Input tensor must not overlap with output tensor.",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Width of input tensor",
                  "direction": "in",
                  "name": "input_x",
                  "type": "const uint16_t"
                },
                {
                  "description": "Height of input tensor",
                  "direction": "in",
                  "name": "input_y",
                  "type": "const uint16_t"
                },
                {
                  "description": "Channels in input tensor",
                  "direction": "in",
                  "name": "input_z",
                  "type": "const uint16_t"
                },
                {
                  "description": "Batch size in input tensor",
                  "direction": "in",
                  "name": "input_w",
                  "type": "const uint16_t"
                },
                {
                  "description": "Pointer to output tensor. Expected to be at least (input_x * input_y * input_z * input_w) + offset_z bytes.",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Channels in output tensor",
                  "direction": "in",
                  "name": "output_z",
                  "type": "const uint16_t"
                },
                {
                  "description": "The offset on the Z axis to start concatenating the input tensor It is user responsibility to provide the correct value",
                  "direction": "in",
                  "name": "offset_z",
                  "type": "const uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_s8_z(\n    const int8_t *input,\n    const uint16_t input_x,\n    const uint16_t input_y,\n    const uint16_t input_z,\n    const uint16_t input_w,\n    int8_t *output,\n    const uint16_t output_z,\n    const uint32_t offset_z\n)",
              "source": {
                "line": 6406,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6406"
              },
              "summary": "int8/uint8 concatenation function to be used for concatenating N-tensors along the Z axis This function should be called for each input tensor to concatenate."
            },
            {
              "description": "int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to concatenate. The argument offset_w will be used to store the input tensor in the correct position in the output tensor\n\ni.e. offset_w = 0 for(i = 0 i < num_input_tensors; ++i) { arm_concatenation_s8_w(&input[i], ..., &output, ..., ..., offset_w) offset_w += input_w[i] }\n\nThis function assumes that the output tensor has:\n\n1. The same width of the input tensor\n2. The same height of the input tensor\n3. The same number o channels of the input tensor\n\nUnless specified otherwise, arguments are mandatory.\n\n:::note\nThis function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation\n\n:::",
              "examples": [],
              "id": "arm_concatenation_s8_w",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_s8_w",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Width of input tensor",
                  "direction": "in",
                  "name": "input_x",
                  "type": "const uint16_t"
                },
                {
                  "description": "Height of input tensor",
                  "direction": "in",
                  "name": "input_y",
                  "type": "const uint16_t"
                },
                {
                  "description": "Channels in input tensor",
                  "direction": "in",
                  "name": "input_z",
                  "type": "const uint16_t"
                },
                {
                  "description": "Batch size in input tensor",
                  "direction": "in",
                  "name": "input_w",
                  "type": "const uint16_t"
                },
                {
                  "description": "Pointer to output tensor. Expected to be at least input_x * input_y * input_z * input_w bytes.",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "The offset on the W axis to start concatenating the input tensor It is user responsibility to provide the correct value",
                  "direction": "in",
                  "name": "offset_w",
                  "type": "const uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_concatenation_s8_w(\n    const int8_t *input,\n    const uint16_t input_x,\n    const uint16_t input_y,\n    const uint16_t input_z,\n    const uint16_t input_w,\n    int8_t *output,\n    const uint32_t offset_w\n)",
              "source": {
                "line": 6449,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6449"
              },
              "summary": "int8/uint8 concatenation function to be used for concatenating N-tensors along the W axis (Batch size) This function should be called for each input tensor to…"
            },
            {
              "description": "int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis\n\n:::note\nThis function, data layout independent, can be used to concatenate either int8 or uint8 tensors because it does not involve any arithmetic operation\n\n:::",
              "examples": [],
              "id": "arm_concatenation_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_s8",
              "params": [
                {
                  "description": "Pointer to input tensors",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *const *"
                },
                {
                  "description": "Number of input tensors",
                  "direction": "in",
                  "name": "inputs_count",
                  "type": "const int32_t"
                },
                {
                  "description": "Dimensions of the input tensors along the target axis",
                  "direction": "in",
                  "name": "input_concat_dims",
                  "type": "const int32_t *"
                },
                {
                  "description": "Target axis to concatenate the input tensors",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor shape",
                  "direction": "in",
                  "name": "output_shape",
                  "type": "const int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_concatenation_s8(\n    const int8_t *const *input_data,\n    const int32_t inputs_count,\n    const int32_t *input_concat_dims,\n    const int32_t axis,\n    int8_t *output_data,\n    const int32_t output_dims,\n    const int32_t *output_shape\n)",
              "source": {
                "line": 6474,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6474"
              },
              "summary": "int8/uint8 concatenation function to be used for concatenating N-tensors along the target axis"
            },
            {
              "description": "int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis\n\n:::note\nThis function, data layout independent, can be used to concatenate either int16 or uint16 tensors because it does not involve any arithmetic operation\n\n:::",
              "examples": [],
              "id": "arm_concatenation_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_s16",
              "params": [
                {
                  "description": "Pointer to input tensors",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *const *"
                },
                {
                  "description": "Number of input tensors",
                  "direction": "in",
                  "name": "inputs_count",
                  "type": "const int32_t"
                },
                {
                  "description": "Dimensions of the input tensors along the target axis",
                  "direction": "in",
                  "name": "input_concat_dims",
                  "type": "const int32_t *"
                },
                {
                  "description": "Target axis to concatenate the input tensors",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor shape",
                  "direction": "in",
                  "name": "output_shape",
                  "type": "const int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_concatenation_s16(\n    const int16_t *const *input_data,\n    const int32_t inputs_count,\n    const int32_t *input_concat_dims,\n    const int32_t axis,\n    int16_t *output_data,\n    const int32_t output_dims,\n    const int32_t *output_shape\n)",
              "source": {
                "line": 6499,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6499"
              },
              "summary": "int16/uint16 concatenation function to be used for concatenating N-tensors along the target axis"
            },
            {
              "description": "int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis\n\n:::note\nThis function, data layout independent, can be used to concatenate either int32 or uint32 tensors because it does not involve any arithmetic operation\n\n:::",
              "examples": [],
              "id": "arm_concatenation_s32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_concatenation_s32",
              "params": [
                {
                  "description": "Pointer to input tensors",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int32_t *const *"
                },
                {
                  "description": "Number of input tensors",
                  "direction": "in",
                  "name": "inputs_count",
                  "type": "const int32_t"
                },
                {
                  "description": "Dimensions of the input tensors along the target axis",
                  "direction": "in",
                  "name": "input_concat_dims",
                  "type": "const int32_t *"
                },
                {
                  "description": "Target axis to concatenate the input tensors",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const int32_t"
                },
                {
                  "description": "Output tensor shape",
                  "direction": "in",
                  "name": "output_shape",
                  "type": "const int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_concatenation_s32(\n    const int32_t *const *input_data,\n    const int32_t inputs_count,\n    const int32_t *input_concat_dims,\n    const int32_t axis,\n    int32_t *output_data,\n    const int32_t output_dims,\n    const int32_t *output_shape\n)",
              "source": {
                "line": 6524,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6524"
              },
              "summary": "int32/uint32 concatenation function to be used for concatenating N-tensors along the target axis"
            },
            {
              "description": "int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis\n\n:::note\nThis function, data layout independent, can be used to split either int8 or uint8 tensors because it does not involve any arithmetic operation.\n\n:::",
              "examples": [],
              "id": "arm_split_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_split_s8",
              "params": [
                {
                  "description": "Pointer to the flattened input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Number of dimensions in input_shape.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const int32_t"
                },
                {
                  "description": "Array of length input_dims describing the shape of input_data.",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Axis along which to split (0 <= axis < input_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of output tensors to produce.",
                  "direction": "in",
                  "name": "num_splits",
                  "type": "const int32_t"
                },
                {
                  "description": "Array of length num_splits giving size of each slice along axis.",
                  "direction": "in",
                  "name": "split_dims",
                  "type": "const int32_t *"
                },
                {
                  "description": "Array of pointers; output_data[i] points to storage for the i-th output tensor.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *const *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR if split_dims sum mismatch."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_split_s8(\n    const int8_t *input_data,\n    const int32_t input_dims,\n    const int32_t *input_shape,\n    const int32_t axis,\n    const int32_t num_splits,\n    const int32_t *split_dims,\n    int8_t *const *output_data\n)",
              "source": {
                "line": 6547,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6547"
              },
              "summary": "int8/uint8 split function to be used for splitting a tensor into multiple tensors along the target axis"
            },
            {
              "description": "int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis\n\n:::note\nThis function, data layout independent, can be used to split either int8 or uint8 tensors because it does not involve any arithmetic operation.\n\n:::",
              "examples": [],
              "id": "arm_split_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_split_s16",
              "params": [
                {
                  "description": "Pointer to the flattened input tensor data.",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Number of dimensions in input_shape.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const int32_t"
                },
                {
                  "description": "Array of length input_dims describing the shape of input_data.",
                  "direction": "in",
                  "name": "input_shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Axis along which to split (0 <= axis < input_dims).",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of output tensors to produce.",
                  "direction": "in",
                  "name": "num_splits",
                  "type": "const int32_t"
                },
                {
                  "description": "Array of length num_splits giving size of each slice along axis.",
                  "direction": "in",
                  "name": "split_dims",
                  "type": "const int32_t *"
                },
                {
                  "description": "Array of pointers; output_data[i] points to storage for the i-th output tensor.",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *const *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS on success, or ARM_CMSIS_NN_ARG_ERROR if split_dims sum mismatch."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_split_s16(\n    const int16_t *input_data,\n    const int32_t input_dims,\n    const int32_t *input_shape,\n    const int32_t axis,\n    const int32_t num_splits,\n    const int32_t *split_dims,\n    int16_t *const *output_data\n)",
              "source": {
                "line": 6570,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6570"
              },
              "summary": "int16/uint16 split function to be used for splitting a tensor into multiple tensors along the target axis"
            },
            {
              "description": "s8 SVDF function with 8 bit state tensor and 8 bit time weights\n\n1. Supported framework: TensorFlow Lite micro",
              "examples": [],
              "id": "arm_svdf_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_s8",
              "params": [
                {
                  "description": "Precomputed per-feature-batch kernel sums, supplied by the caller. This is an input the function only reads, not scratch it fills: an allocated but unfilled buffer yields wrong output while still returning ARM_CMSIS_NN_SUCCESS. Mandatory on builds with the MVE extension (ARM_MATH_MVEI), where a NULL ctx->buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. Unused on every other build, where ctx->buf may be NULL. Sized by arm_svdf_s8_get_buffer_size(weights_feature_dims): weights_feature_dims->n * sizeof(int32_t) where the sums are used, 0 otherwise. Note this is weights_feature_dims->n, not a filter_dims->c - do not size this buffer with `arm_fully_connected_s8_get_buffer_size()`, which reads a different field and under-allocates. Fill it with arm_vector_sum_s8(ctx->buf, input_dims->h, weights_feature_dims->n, weights_feature_data, -svdf_params->input_offset, 0, NULL) so that entry j holds -input_offset * sum(weights_feature row j). The contents depend only on weights_feature_data and svdf_params->input_offset, so they may be computed once at load time and reused across calls until one of those changes. The buffer is specific to one layer's weights and cannot be shared between layers. Do NOT clear this buffer between calls: zeroing it is indistinguishable from leaving it unfilled, and produces the silently wrong output described above. If it must be cleared for security reasons, clear it after the last call that uses it, and refill it before any further call.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Scratch buffer written by this function, holding one int32_t accumulator per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_s8_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(int32_t) bytes, the same figure on every build target. This function does not read input_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "input_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Scratch buffer written by this function, holding one int32_t accumulator per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_s8_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t) bytes, truncating division, the same figure on every build target. This function does not read output_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "SVDF Parameters Range of svdf_params->input_offset : [-128, 127] Range of svdf_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "svdf_params",
                  "type": "const cmsis_nn_svdf_params *"
                },
                {
                  "description": "Input quantization parameters",
                  "direction": "in",
                  "name": "input_quant_params",
                  "type": "const cmsis_nn_per_tensor_quant_params *"
                },
                {
                  "description": "Output quantization parameters",
                  "direction": "in",
                  "name": "output_quant_params",
                  "type": "const cmsis_nn_per_tensor_quant_params *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "State tensor dimensions",
                  "direction": "in",
                  "name": "state_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to state tensor",
                  "direction": "inout",
                  "name": "state_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Weights (feature) tensor dimensions",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the weights (feature) tensor",
                  "direction": "in",
                  "name": "weights_feature_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Weights (time) tensor dimensions",
                  "direction": "in",
                  "name": "weights_time_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the weights (time) tensor",
                  "direction": "in",
                  "name": "weights_time_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias tensor dimensions",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to bias tensor",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns either `ARM_CMSIS_NN_ARG_ERROR` if argument constraints fail. or, `ARM_CMSIS_NN_SUCCESS` on successful completion."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_svdf_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *input_ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_svdf_params *svdf_params,\n    const cmsis_nn_per_tensor_quant_params *input_quant_params,\n    const cmsis_nn_per_tensor_quant_params *output_quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *state_dims,\n    int8_t *state_data,\n    const cmsis_nn_dims *weights_feature_dims,\n    const int8_t *weights_feature_data,\n    const cmsis_nn_dims *weights_time_dims,\n    const int8_t *weights_time_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 6662,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6662"
              },
              "summary": "s8 SVDF function with 8 bit state tensor and 8 bit time weights"
            },
            {
              "description": "s8 SVDF function with 16 bit state tensor and 16 bit time weights\n\n1. Supported framework: TensorFlow Lite micro",
              "examples": [],
              "id": "arm_svdf_state_s16_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_state_s16_s8",
              "params": [
                {
                  "description": "Scratch buffer written by this function, holding one int32_t accumulator per (input batch, feature batch). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_state_s16_s8_input_ctx_get_buffer_size(input_dims, weights_feature_dims): input_dims->n * weights_feature_dims->n * sizeof(int32_t) bytes, the same figure on every build target. Note the accumulators are int32_t even though the state tensor is int16_t - this buffer does not shrink with the state width. This function does not read input_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "input_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Scratch buffer written by this function, holding one int32_t accumulator per (input batch, output unit). Written before it is read, so its contents on entry do not matter, but it is written on EVERY build, not only under MVE. Mandatory: a NULL buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR on every build. Sized by arm_svdf_state_s16_s8_output_ctx_get_buffer_size(svdf_params, input_dims, weights_feature_dims): input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t) bytes, truncating division, the same figure on every build target. This function does not read output_ctx->size, so an undersized buffer is not diagnosed: query the sizer above and honour it. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "output_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "SVDF Parameters Range of svdf_params->input_offset : [-128, 127] Range of svdf_params->output_offset : [-128, 127]",
                  "direction": "in",
                  "name": "svdf_params",
                  "type": "const cmsis_nn_svdf_params *"
                },
                {
                  "description": "Input quantization parameters",
                  "direction": "in",
                  "name": "input_quant_params",
                  "type": "const cmsis_nn_per_tensor_quant_params *"
                },
                {
                  "description": "Output quantization parameters",
                  "direction": "in",
                  "name": "output_quant_params",
                  "type": "const cmsis_nn_per_tensor_quant_params *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "State tensor dimensions",
                  "direction": "in",
                  "name": "state_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to state tensor",
                  "direction": "inout",
                  "name": "state_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Weights (feature) tensor dimensions",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the weights (feature) tensor",
                  "direction": "in",
                  "name": "weights_feature_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Weights (time) tensor dimensions",
                  "direction": "in",
                  "name": "weights_time_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the weights (time) tensor",
                  "direction": "in",
                  "name": "weights_time_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Bias tensor dimensions",
                  "direction": "in",
                  "name": "bias_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to bias tensor",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_svdf_state_s16_s8(\n    const cmsis_nn_context *input_ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_svdf_params *svdf_params,\n    const cmsis_nn_per_tensor_quant_params *input_quant_params,\n    const cmsis_nn_per_tensor_quant_params *output_quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input_data,\n    const cmsis_nn_dims *state_dims,\n    int16_t *state_data,\n    const cmsis_nn_dims *weights_feature_dims,\n    const int8_t *weights_feature_data,\n    const cmsis_nn_dims *weights_time_dims,\n    const int16_t *weights_time_data,\n    const cmsis_nn_dims *bias_dims,\n    const int32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output_data\n)",
              "source": {
                "line": 6734,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6734"
              },
              "summary": "s8 SVDF function with 16 bit state tensor and 16 bit time weights"
            },
            {
              "description": "Get size of the kernel-sum buffer required by `arm_svdf_s8()`.\n\nFor a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. `arm_svdf_s8()` has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other `cmsis_nn_dims` of that call can size it - in particular `arm_fully_connected_s8_get_buffer_size()` reads a different field and under-allocates. See `arm_svdf_s8()` for the buffer's layout, how to fill it and when it may be reused.",
              "examples": [],
              "id": "arm_svdf_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_s8_get_buffer_size",
              "params": [
                {
                  "description": "dimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_svdf_s8_get_buffer_size(const cmsis_nn_dims *weights_feature_dims)",
              "source": {
                "line": 6767,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6767"
              },
              "summary": "Get size of the kernel-sum buffer required by armsvdfs8()."
            },
            {
              "description": "Get size of the kernel-sum buffer required by `arm_svdf_s8()` for processors with DSP extension.\n\nFor a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. `arm_svdf_s8()` has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other `cmsis_nn_dims` of that call can size it - in particular `arm_fully_connected_s8_get_buffer_size()` reads a different field and under-allocates. See `arm_svdf_s8()` for the buffer's layout, how to fill it and when it may be reused.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_svdf_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAlso validates dims like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_svdf_s8_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_s8_get_buffer_size_dsp",
              "params": [
                {
                  "description": "dimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_svdf_s8_get_buffer_size_dsp(const cmsis_nn_dims *weights_feature_dims)",
              "source": {
                "line": 6778,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6778"
              },
              "summary": "Get size of the kernel-sum buffer required by armsvdfs8() for processors with DSP extension."
            },
            {
              "description": "Get size of the kernel-sum buffer required by `arm_svdf_s8()` for Arm(R) Helium Architecture case.\n\nFor a valid (non-negative, in-range) weights_feature_dims->n, returns weights_feature_dims->n * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid weights_feature_dims->n, returns -1 on every build target. `arm_svdf_s8()` has no filter_dims argument, and the buffer is indexed by weights_feature_dims->n, so no other `cmsis_nn_dims` of that call can size it - in particular `arm_fully_connected_s8_get_buffer_size()` reads a different field and under-allocates. See `arm_svdf_s8()` for the buffer's layout, how to fill it and when it may be reused.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_svdf_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAlso validates dims like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_svdf_s8_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_s8_get_buffer_size_mve",
              "params": [
                {
                  "description": "dimensions of the weights (feature) tensor, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if weights_feature_dims->n is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_svdf_s8_get_buffer_size_mve(const cmsis_nn_dims *weights_feature_dims)",
              "source": {
                "line": 6789,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6789"
              },
              "summary": "Get size of the kernel-sum buffer required by armsvdfs8() for Arm(R) Helium Architecture case."
            },
            {
              "description": "Get size of the input_ctx staging buffer required by `arm_svdf_s8()`.\n\nReturns input_dims->n * weights_feature_dims->n * sizeof(int32_t). Unlike `arm_svdf_s8_get_buffer_size()`, this figure does not vary by build target: `arm_svdf_s8()` stages this buffer on every build, not only under MVE, so there is no _dsp / _mve pair to choose between and the validation runs on every target.\n\n:::note\nThis is a different buffer from the one `arm_svdf_s8_get_buffer_size()` describes. That one sizes the read-only kernel sums passed as ctx; this one sizes the scratch passed as input_ctx.\n\n:::\n\n:::note\n0 is a valid return for a degenerate shape (input_dims->n == 0). Unlike the general rule in README.md, a 0 here does NOT mean you may pass { NULL, 0 }: `arm_svdf_s8()` rejects a NULL input_ctx->buf with ARM_CMSIS_NN_ARG_ERROR regardless of the size. -1 is used only for an out-of-range or NULL argument.\n\n:::",
              "examples": [],
              "id": "arm_svdf_s8_input_ctx_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_s8_input_ctx_get_buffer_size",
              "params": [
                {
                  "description": "Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_svdf_s8_input_ctx_get_buffer_size(\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *weights_feature_dims\n)",
              "source": {
                "line": 6812,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6812"
              },
              "summary": "Get size of the inputctx staging buffer required by armsvdfs8()."
            },
            {
              "description": "Get size of the output_ctx staging buffer required by `arm_svdf_s8()`.\n\nReturns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t). The division truncates, matching the kernel's own unit count. As with `arm_svdf_s8_input_ctx_get_buffer_size()`, the figure is the same on every build target and the validation runs on every target.\n\n:::note\nSame degenerate-0 contract as `arm_svdf_s8_input_ctx_get_buffer_size()`, including that a 0 does not license passing { NULL, 0 }. A rank greater than weights_feature_dims->n truncates the unit count to 0 and so returns 0.\n\n:::\n\n:::note\n`arm_svdf_s8()` narrows svdf_params->rank to int16_t before dividing by it, so a rank outside int16_t range would make this query and the kernel disagree - 65538 narrows to 2. The kernel can then write unboundedly more than the untruncated formula reports, because that formula truncates to 0 whenever weights_feature_dims->n < 65538: at weights_feature_dims->n = 100 it would report 0 bytes while the kernel writes 50 units, i.e. 200 bytes. Such a rank returns -1 rather than a number the kernel will not honour. Ranks that survive the int16_t round trip, that is within [-32768, 32767], are unaffected; this library does not otherwise constrain svdf_params->rank.\n\n:::",
              "examples": [],
              "id": "arm_svdf_s8_output_ctx_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_s8_output_ctx_get_buffer_size",
              "params": [
                {
                  "description": "SVDF parameters; only svdf_params->rank is read",
                  "direction": "in",
                  "name": "svdf_params",
                  "type": "const cmsis_nn_svdf_params *"
                },
                {
                  "description": "Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any pointer is NULL, if svdf_params->rank is zero, negative or outside int16_t range, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_svdf_s8_output_ctx_get_buffer_size(\n    const cmsis_nn_svdf_params *svdf_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *weights_feature_dims\n)",
              "source": {
                "line": 6842,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6842"
              },
              "summary": "Get size of the outputctx staging buffer required by armsvdfs8()."
            },
            {
              "description": "Get size of the input_ctx staging buffer required by `arm_svdf_state_s16_s8()`.\n\nReturns input_dims->n * weights_feature_dims->n * sizeof(int32_t). Unlike `arm_svdf_s8_get_buffer_size()`, this figure does not vary by build target: `arm_svdf_s8()` stages this buffer on every build, not only under MVE, so there is no _dsp / _mve pair to choose between and the validation runs on every target.\n\n:::note\nThis is a different buffer from the one `arm_svdf_s8_get_buffer_size()` describes. That one sizes the read-only kernel sums passed as ctx; this one sizes the scratch passed as input_ctx.\n\n:::\n\n:::note\n0 is a valid return for a degenerate shape (input_dims->n == 0). Unlike the general rule in README.md, a 0 here does NOT mean you may pass { NULL, 0 }: `arm_svdf_s8()` rejects a NULL input_ctx->buf with ARM_CMSIS_NN_ARG_ERROR regardless of the size. -1 is used only for an out-of-range or NULL argument.\n\n:::\n\nReturns input_dims->n * weights_feature_dims->n * sizeof(int32_t) - the same figure as `arm_svdf_s8_input_ctx_get_buffer_size()` for the same shape. The accumulators are int32_t even though `arm_svdf_state_s16_s8()` carries an int16_t state tensor, so this buffer does not shrink with the state width.",
              "examples": [],
              "id": "arm_svdf_state_s16_s8_input_ctx_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_state_s16_s8_input_ctx_get_buffer_size",
              "params": [
                {
                  "description": "Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if either pointer is NULL, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_svdf_state_s16_s8_input_ctx_get_buffer_size(\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *weights_feature_dims\n)",
              "source": {
                "line": 6855,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6855"
              },
              "summary": "Get size of the inputctx staging buffer required by armsvdfstates16s8()."
            },
            {
              "description": "Get size of the output_ctx staging buffer required by `arm_svdf_state_s16_s8()`.\n\nReturns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t). The division truncates, matching the kernel's own unit count. As with `arm_svdf_s8_input_ctx_get_buffer_size()`, the figure is the same on every build target and the validation runs on every target.\n\n:::note\nSame degenerate-0 contract as `arm_svdf_s8_input_ctx_get_buffer_size()`, including that a 0 does not license passing { NULL, 0 }. A rank greater than weights_feature_dims->n truncates the unit count to 0 and so returns 0.\n\n:::\n\n:::note\n`arm_svdf_s8()` narrows svdf_params->rank to int16_t before dividing by it, so a rank outside int16_t range would make this query and the kernel disagree - 65538 narrows to 2. The kernel can then write unboundedly more than the untruncated formula reports, because that formula truncates to 0 whenever weights_feature_dims->n < 65538: at weights_feature_dims->n = 100 it would report 0 bytes while the kernel writes 50 units, i.e. 200 bytes. Such a rank returns -1 rather than a number the kernel will not honour. Ranks that survive the int16_t round trip, that is within [-32768, 32767], are unaffected; this library does not otherwise constrain svdf_params->rank.\n\n:::\n\nReturns input_dims->n * (weights_feature_dims->n / svdf_params->rank) * sizeof(int32_t), truncating division - the same figure as `arm_svdf_s8_output_ctx_get_buffer_size()` for the same shape.",
              "examples": [],
              "id": "arm_svdf_state_s16_s8_output_ctx_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_svdf_state_s16_s8_output_ctx_get_buffer_size",
              "params": [
                {
                  "description": "SVDF parameters; only svdf_params->rank is read",
                  "direction": "in",
                  "name": "svdf_params",
                  "type": "const cmsis_nn_svdf_params *"
                },
                {
                  "description": "Input tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Weights (feature) tensor dimensions, i.e. the same `cmsis_nn_dims` passed to `arm_svdf_s8()`",
                  "direction": "in",
                  "name": "weights_feature_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any pointer is NULL, if svdf_params->rank is zero, negative or outside int16_t range, if input_dims->n or weights_feature_dims->n is negative, or if the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_svdf_state_s16_s8_output_ctx_get_buffer_size(\n    const cmsis_nn_svdf_params *svdf_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *weights_feature_dims\n)",
              "source": {
                "line": 6865,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6865"
              },
              "summary": "Get size of the outputctx staging buffer required by armsvdfstates16s8()."
            },
            {
              "description": "LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_lstm_unidirectional_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_s8",
              "params": [
                {
                  "description": "Pointer to input data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to output data",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Struct containing all information about the lstm operator, see arm_nn_types.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params *"
                },
                {
                  "description": "Struct containing pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. Size temp1 with `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` and temp2 with `arm_lstm_unidirectional_s8_temp2_get_buffer_size()` - both hold int16_t gate vectors even though the layer datatype is s8, so sizing them in s8 elements under-allocates by half.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_lstm_context *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_lstm_unidirectional_s8(\n    const int8_t *input,\n    int8_t *output,\n    const cmsis_nn_lstm_params *params,\n    cmsis_nn_lstm_context *buffers\n)",
              "source": {
                "line": 6892,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6892"
              },
              "summary": "LSTM unidirectional function with 8 bit input and output and 16 bit gate output, 32 bit bias."
            },
            {
              "description": "LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_lstm_unidirectional_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_s16",
              "params": [
                {
                  "description": "Pointer to input data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to output data",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Struct containing all information about the lstm operator, see arm_nn_types.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params *"
                },
                {
                  "description": "Struct containing pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types. Size temp1 with `arm_lstm_unidirectional_s16_temp1_get_buffer_size()` and temp2 with `arm_lstm_unidirectional_s16_temp2_get_buffer_size()`.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_lstm_context *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_lstm_unidirectional_s16(\n    const int16_t *input,\n    int16_t *output,\n    const cmsis_nn_lstm_params *params,\n    cmsis_nn_lstm_context *buffers\n)",
              "source": {
                "line": 6914,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6914"
              },
              "summary": "LSTM unidirectional function with 16 bit input and output and 16 bit gate output, 64 bit bias."
            },
            {
              "description": "Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_s8()`.\n\n:::note\ntime_steps does not enter the requirement: the buffer is reused by every step. A layer with time_steps == 0 runs no step and never dereferences the buffer, but the query still reports the per-step figure rather than 0.\n\n:::\n\n:::note\n0 is only returned for a degenerate shape (batch_size or hidden_size of 0), for which `arm_lstm_unidirectional_s8()` makes no scratch access, so { NULL } is acceptable there per the README.md buffer convention. There is no runtime enforcement: the kernel does not range-check the buffers, and an undersized allocation is written past on every build target.\n\n:::",
              "examples": [],
              "id": "arm_lstm_unidirectional_s8_temp1_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_s8_temp1_get_buffer_size",
              "params": [
                {
                  "description": "LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read.",
                  "direction": "in",
                  "name": "lstm_params",
                  "type": "const cmsis_nn_lstm_params *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target."
                }
              ],
              "signature": "int32_t arm_lstm_unidirectional_s8_temp1_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)",
              "source": {
                "line": 6940,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6940"
              },
              "summary": "Get size of the temp1 scratch buffer required by armlstmunidirectionals8()."
            },
            {
              "description": "Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_s8()`.\n\n:::note\ntime_steps does not enter the requirement: the buffer is reused by every step. A layer with time_steps == 0 runs no step and never dereferences the buffer, but the query still reports the per-step figure rather than 0.\n\n:::\n\n:::note\n0 is only returned for a degenerate shape (batch_size or hidden_size of 0), for which `arm_lstm_unidirectional_s8()` makes no scratch access, so { NULL } is acceptable there per the README.md buffer convention. There is no runtime enforcement: the kernel does not range-check the buffers, and an undersized allocation is written past on every build target.\n\n:::",
              "examples": [],
              "id": "arm_lstm_unidirectional_s8_temp2_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_s8_temp2_get_buffer_size",
              "params": [
                {
                  "description": "LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read.",
                  "direction": "in",
                  "name": "lstm_params",
                  "type": "const cmsis_nn_lstm_params *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target."
                },
                {
                  "description": "Required buffer size in bytes: the same figure as `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` for the same params. temp2 stages the cell-gate vector and the tanh(cell_state) vector, both of the same extent as the gate vectors staged in temp1."
                }
              ],
              "signature": "int32_t arm_lstm_unidirectional_s8_temp2_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)",
              "source": {
                "line": 6950,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6950"
              },
              "summary": "Get size of the temp2 scratch buffer required by armlstmunidirectionals8()."
            },
            {
              "description": "Get size of the temp1 scratch buffer required by `arm_lstm_unidirectional_s16()`.\n\n:::note\ntime_steps does not enter the requirement: the buffer is reused by every step. A layer with time_steps == 0 runs no step and never dereferences the buffer, but the query still reports the per-step figure rather than 0.\n\n:::\n\n:::note\n0 is only returned for a degenerate shape (batch_size or hidden_size of 0), for which `arm_lstm_unidirectional_s8()` makes no scratch access, so { NULL } is acceptable there per the README.md buffer convention. There is no runtime enforcement: the kernel does not range-check the buffers, and an undersized allocation is written past on every build target.\n\n:::",
              "examples": [],
              "id": "arm_lstm_unidirectional_s16_temp1_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_s16_temp1_get_buffer_size",
              "params": [
                {
                  "description": "LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read.",
                  "direction": "in",
                  "name": "lstm_params",
                  "type": "const cmsis_nn_lstm_params *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target."
                },
                {
                  "description": "Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t) - the same figure as `arm_lstm_unidirectional_s8_temp1_get_buffer_size()` for the same params, since both layer datatypes stage int16_t gate vectors."
                }
              ],
              "signature": "int32_t arm_lstm_unidirectional_s16_temp1_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)",
              "source": {
                "line": 6961,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6961"
              },
              "summary": "Get size of the temp1 scratch buffer required by armlstmunidirectionals16()."
            },
            {
              "description": "Get size of the temp2 scratch buffer required by `arm_lstm_unidirectional_s16()`.\n\n:::note\ntime_steps does not enter the requirement: the buffer is reused by every step. A layer with time_steps == 0 runs no step and never dereferences the buffer, but the query still reports the per-step figure rather than 0.\n\n:::\n\n:::note\n0 is only returned for a degenerate shape (batch_size or hidden_size of 0), for which `arm_lstm_unidirectional_s8()` makes no scratch access, so { NULL } is acceptable there per the README.md buffer convention. There is no runtime enforcement: the kernel does not range-check the buffers, and an undersized allocation is written past on every build target.\n\n:::",
              "examples": [],
              "id": "arm_lstm_unidirectional_s16_temp2_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_lstm_unidirectional_s16_temp2_get_buffer_size",
              "params": [
                {
                  "description": "LSTM operator parameters, i.e. the same `cmsis_nn_lstm_params` passed to `arm_lstm_unidirectional_s8()`. Only time_major, batch_size and hidden_size are read.",
                  "direction": "in",
                  "name": "lstm_params",
                  "type": "const cmsis_nn_lstm_params *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Required buffer size in bytes: (time_major != 0 ? batch_size : 1) * hidden_size * sizeof(int16_t). The elements are int16_t gate outputs even though the layer datatype is s8. batch_size enters only for a time-major layer because the batch-major wrapper always re-invokes the step kernel one batch at a time. Returns -1 if lstm_params is NULL, if batch_size or hidden_size is negative, or if the product would not fit in an int32_t. The figure and the range checks are the same on every build target."
                },
                {
                  "description": "Required buffer size in bytes: the same figure as `arm_lstm_unidirectional_s16_temp1_get_buffer_size()` for the same params."
                }
              ],
              "signature": "int32_t arm_lstm_unidirectional_s16_temp2_get_buffer_size(const cmsis_nn_lstm_params *lstm_params)",
              "source": {
                "line": 6970,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L6970"
              },
              "summary": "Get size of the temp2 scratch buffer required by armlstmunidirectionals16()."
            },
            {
              "description": "Batch matmul function with 8 bit input and output.\n\n1. Supported framework: TensorFlow Lite Micro\n2. Performs row * row matrix multiplication with the RHS transposed.",
              "examples": [],
              "id": "arm_batch_matmul_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_matmul_s8",
              "params": [
                {
                  "description": "Temporary scratch buffer for the per-row kernel sums of the RHS. Mandatory on builds with the MVE extension (ARM_MATH_MVEI), where a NULL ctx->buf is diagnosed with ARM_CMSIS_NN_ARG_ERROR. Unused on every other build, where ctx->buf may be NULL. Sized by arm_batch_matmul_s8_get_buffer_size(input_rhs_dims) - pass the same input_rhs_dims given below. That is input_rhs_dims->w * sizeof(int32_t) where the sums are used, 0 otherwise. Do not size this buffer with `arm_fully_connected_s8_get_buffer_size()`: it reads a different field, and an allocation short of input_rhs_dims->w words is written past its end. The function fills the buffer itself before each use, so the caller does not need to initialize it. ctx->buf must be aligned to sizeof(int32_t). If ctx->size is non-zero it is validated against the requirement and a buffer too small is rejected with ARM_CMSIS_NN_ARG_ERROR; a ctx->size of 0 skips that check. A negative input_rhs_dims->w, or one large enough that the required size exceeds INT32_MAX, is rejected with ARM_CMSIS_NN_ARG_ERROR regardless of ctx->size. The caller is expected to clear the buffer, if applicable, for security reasons.",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Batch matmul Parameters Adjoint flags are currently unused and do not transpose either input; callers must supply the tensors in the layouts described below.",
                  "direction": "in",
                  "name": "bmm_params",
                  "type": "const cmsis_nn_bmm_params *"
                },
                {
                  "description": "Quantization parameters",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_tensor_quant_params *"
                },
                {
                  "description": "Input lhs tensor dimensions. This s8 function treats w as the row count and c as the inner dimension. This differs from `arm_batch_matmul_f32()`, so its dimension mapping must not be reused here.",
                  "direction": "in",
                  "name": "input_lhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input rhs tensor dimensions. The RHS must already be transposed, with w as its row count and c equal to input_lhs_dims->c.",
                  "direction": "in",
                  "name": "input_rhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to transposed input tensor",
                  "direction": "in",
                  "name": "input_rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the following:\n\n- `ARM_CMSIS_NN_ARG_ERROR` if an MVE build receives an invalid context, a negative or unrepresentable RHS row count, or a declared context size below the requirement.\n- `ARM_CMSIS_NN_SUCCESS` on success."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_batch_matmul_s8(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_bmm_params *bmm_params,\n    const cmsis_nn_per_tensor_quant_params *quant_params,\n    const cmsis_nn_dims *input_lhs_dims,\n    const int8_t *input_lhs,\n    const cmsis_nn_dims *input_rhs_dims,\n    const int8_t *input_rhs,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output\n)",
              "source": {
                "line": 7017,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7017"
              },
              "summary": "Batch matmul function with 8 bit input and output."
            },
            {
              "description": "Batch matmul function with 16 bit input and output.\n\n1. Supported framework: TensorFlow Lite Micro\n2. Performs row * row matrix multiplication with the RHS transposed.",
              "examples": [],
              "id": "arm_batch_matmul_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_matmul_s16",
              "params": [
                {
                  "description": "Unused: this function requires no scratch buffer and does not read or write ctx on any build, so ctx->buf may be NULL. Retained for signature compatibility with `arm_batch_matmul_s8()`. There is deliberately no arm_batch_matmul_s16_get_buffer_size(); in particular `arm_fully_connected_s8_get_buffer_size()` is not the sizer for this argument. If a real buffer is passed, the caller is expected to clear it, if applicable, for security reasons.",
                  "direction": "in",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Batch matmul Parameters Adjoint flags are currently unused.",
                  "direction": "in",
                  "name": "bmm_params",
                  "type": "const cmsis_nn_bmm_params *"
                },
                {
                  "description": "Quantization parameters",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_tensor_quant_params *"
                },
                {
                  "description": "Input lhs tensor dimensions. This should be NHWC where LHS.C = RHS.C",
                  "direction": "in",
                  "name": "input_lhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_lhs",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input lhs tensor dimensions. This is expected to be transposed so should be NHWC where LHS.C = RHS.C",
                  "direction": "in",
                  "name": "input_rhs_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to transposed input tensor",
                  "direction": "in",
                  "name": "input_rhs",
                  "type": "const int16_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_batch_matmul_s16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_bmm_params *bmm_params,\n    const cmsis_nn_per_tensor_quant_params *quant_params,\n    const cmsis_nn_dims *input_lhs_dims,\n    const int16_t *input_lhs,\n    const cmsis_nn_dims *input_rhs_dims,\n    const int16_t *input_rhs,\n    const cmsis_nn_dims *output_dims,\n    int16_t *output\n)",
              "source": {
                "line": 7057,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7057"
              },
              "summary": "Batch matmul function with 16 bit input and output."
            },
            {
              "description": "Get size of the scratch buffer required by `arm_batch_matmul_s8()`.\n\nFor a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with `arm_fully_connected_s8_get_buffer_size()`, which reads .c) writes past the allocation whenever the rhs has more rows than columns. `arm_batch_matmul_s16()` needs no scratch buffer and so has no corresponding sizer.",
              "examples": [],
              "id": "arm_batch_matmul_s8_get_buffer_size",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_matmul_s8_get_buffer_size",
              "params": [
                {
                  "description": "dimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()`",
                  "direction": "in",
                  "name": "input_rhs_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_batch_matmul_s8_get_buffer_size(const cmsis_nn_dims *input_rhs_dims)",
              "source": {
                "line": 7082,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7082"
              },
              "summary": "Get size of the scratch buffer required by armbatchmatmuls8()."
            },
            {
              "description": "Get size of the scratch buffer required by `arm_batch_matmul_s8()` for processors with DSP extension.\n\nFor a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with `arm_fully_connected_s8_get_buffer_size()`, which reads .c) writes past the allocation whenever the rhs has more rows than columns. `arm_batch_matmul_s16()` needs no scratch buffer and so has no corresponding sizer.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_batch_matmul_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAlso validates dims like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_batch_matmul_s8_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_matmul_s8_get_buffer_size_dsp",
              "params": [
                {
                  "description": "dimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()`",
                  "direction": "in",
                  "name": "input_rhs_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_batch_matmul_s8_get_buffer_size_dsp(const cmsis_nn_dims *input_rhs_dims)",
              "source": {
                "line": 7093,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7093"
              },
              "summary": "Get size of the scratch buffer required by armbatchmatmuls8() for processors with DSP extension."
            },
            {
              "description": "Get size of the scratch buffer required by `arm_batch_matmul_s8()` for Arm(R) Helium Architecture case.\n\nFor a valid (non-negative, in-range) input_rhs_dims->w, returns input_rhs_dims->w * sizeof(int32_t) on builds with the MVE extension and 0 elsewhere. For an invalid input_rhs_dims->w, returns -1 on every build target. input_rhs_dims->w is the rhs row count, which is what the kernel-sum buffer is indexed by; sizing this buffer from any other dims (in particular with `arm_fully_connected_s8_get_buffer_size()`, which reads .c) writes past the allocation whenever the rhs has more rows than columns. `arm_batch_matmul_s16()` needs no scratch buffer and so has no corresponding sizer.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_batch_matmul_s8_get_buffer_size()`.\n\n:::\n\n:::note\nAlso validates dims like the top-level dispatcher, returning -1 for invalid values.\n\n:::",
              "examples": [],
              "id": "arm_batch_matmul_s8_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_batch_matmul_s8_get_buffer_size_mve",
              "params": [
                {
                  "description": "dimensions of the (transposed) rhs tensor, i.e. the same `cmsis_nn_dims` passed to `arm_batch_matmul_s8()`",
                  "direction": "in",
                  "name": "input_rhs_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if input_rhs_dims->w is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_batch_matmul_s8_get_buffer_size_mve(const cmsis_nn_dims *input_rhs_dims)",
              "source": {
                "line": 7104,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7104"
              },
              "summary": "Get size of the scratch buffer required by armbatchmatmuls8() for Arm(R) Helium Architecture case."
            },
            {
              "description": "Expands the size of the input by adding constant values before and after the data, in all dimensions.",
              "examples": [],
              "id": "arm_pad_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_pad_s8",
              "params": [
                {
                  "description": "Pointer to input data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to output data",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Value to pad with",
                  "direction": "in",
                  "name": "pad_value",
                  "type": "const int8_t"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_size",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Padding to apply before data in each dimension",
                  "direction": "in",
                  "name": "pre_pad",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Padding to apply after data in each dimension",
                  "direction": "in",
                  "name": "post_pad",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_pad_s8(\n    const int8_t *input,\n    int8_t *output,\n    const int8_t pad_value,\n    const cmsis_nn_dims *input_size,\n    const cmsis_nn_dims *pre_pad,\n    const cmsis_nn_dims *post_pad\n)",
              "source": {
                "line": 7124,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7124"
              },
              "summary": "Expands the size of the input by adding constant values before and after the data, in all dimensions."
            },
            {
              "description": "Expands the size of the input by adding constant values before and after the data, in all dimensions.",
              "examples": [],
              "id": "arm_pad_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_pad_s16",
              "params": [
                {
                  "description": "Pointer to input data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to output data",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Value to pad with",
                  "direction": "in",
                  "name": "pad_value",
                  "type": "const int16_t"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_size",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Padding to apply before data in each dimension",
                  "direction": "in",
                  "name": "pre_pad",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Padding to apply after data in each dimension",
                  "direction": "in",
                  "name": "post_pad",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_pad_s16(\n    const int16_t *input,\n    int16_t *output,\n    const int16_t pad_value,\n    const cmsis_nn_dims *input_size,\n    const cmsis_nn_dims *pre_pad,\n    const cmsis_nn_dims *post_pad\n)",
              "source": {
                "line": 7144,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7144"
              },
              "summary": "Expands the size of the input by adding constant values before and after the data, in all dimensions."
            },
            {
              "description": "Computes the mean of the input tensor along the specified axis. The output multipler and shift must have the output count folded into them.",
              "examples": [],
              "id": "arm_mean_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_mean_s8",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Axis dimensions to compute mean over",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output quantization multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Output quantization shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_mean_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t input_offset,\n    const cmsis_nn_dims *axis_dims,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift\n)",
              "source": {
                "line": 7173,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7173"
              },
              "summary": "Computes the mean of the input tensor along the specified axis."
            },
            {
              "description": "Computes the mean of the input tensor along the specified axis. The output multipler and shift must have the output count folded into them.",
              "examples": [],
              "id": "arm_mean_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_mean_s16",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input offset",
                  "direction": "in",
                  "name": "input_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Axis dimensions to compute mean over",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output quantization multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Output quantization shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_mean_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t input_offset,\n    const cmsis_nn_dims *axis_dims,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift\n)",
              "source": {
                "line": 7200,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7200"
              },
              "summary": "Computes the mean of the input tensor along the specified axis."
            },
            {
              "description": "Compute ArgMax indices of an s8 tensor along a specific axis.",
              "examples": [],
              "id": "arm_argmax_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_argmax_s8",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions (NHWC layout)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Reduction axis in range [0, 3]",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to output indices (int32_t)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_argmax_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t axis,\n    int32_t *output_data\n)",
              "source": {
                "line": 7221,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7221"
              },
              "summary": "Compute ArgMax indices of an s8 tensor along a specific axis."
            },
            {
              "description": "Compute ArgMin indices of an s8 tensor along a specific axis.",
              "examples": [],
              "id": "arm_argmin_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_argmin_s8",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions (NHWC layout)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Reduction axis in range [0, 3]",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to output indices (int32_t)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_argmin_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t axis,\n    int32_t *output_data\n)",
              "source": {
                "line": 7234,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7234"
              },
              "summary": "Compute ArgMin indices of an s8 tensor along a specific axis."
            },
            {
              "description": "Compute ArgMax indices of an s16 tensor along a specific axis.",
              "examples": [],
              "id": "arm_argmax_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_argmax_s16",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions (NHWC layout)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Reduction axis in range [0, 3]",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to output indices (int32_t)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_argmax_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t axis,\n    int32_t *output_data\n)",
              "source": {
                "line": 7247,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7247"
              },
              "summary": "Compute ArgMax indices of an s16 tensor along a specific axis."
            },
            {
              "description": "Compute ArgMin indices of an s16 tensor along a specific axis.",
              "examples": [],
              "id": "arm_argmin_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_argmin_s16",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions (NHWC layout)",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Reduction axis in range [0, 3]",
                  "direction": "in",
                  "name": "axis",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to output indices (int32_t)",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` for a valid axis, otherwise `ARM_CMSIS_NN_ARG_ERROR`."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_argmin_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t axis,\n    int32_t *output_data\n)",
              "source": {
                "line": 7260,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7260"
              },
              "summary": "Compute ArgMin indices of an s16 tensor along a specific axis."
            },
            {
              "description": "Computes the max of the input tensor along the specified axis.",
              "examples": [],
              "id": "arm_reduce_max_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_max_s8",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Axis dimensions to compute mean over",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_max_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 7274,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7274"
              },
              "summary": "Computes the max of the input tensor along the specified axis."
            },
            {
              "description": "Computes the max of the input tensor along the specified axis.",
              "examples": [],
              "id": "arm_reduce_max_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_max_s16",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Axis dimensions to compute mean over",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_max_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 7292,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7292"
              },
              "summary": "Computes the max of the input tensor along the specified axis."
            },
            {
              "description": "Computes the min of the input tensor along the specified axis.",
              "examples": [],
              "id": "arm_reduce_min_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_min_s8",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Axis dimensions to compute mean over",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_min_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 7310,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7310"
              },
              "summary": "Computes the min of the input tensor along the specified axis."
            },
            {
              "description": "Computes the min of the input tensor along the specified axis.",
              "examples": [],
              "id": "arm_reduce_min_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_min_s16",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Axis dimensions to compute mean over",
                  "direction": "in",
                  "name": "axis_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reduce_min_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *axis_dims,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 7328,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7328"
              },
              "summary": "Computes the min of the input tensor along the specified axis."
            },
            {
              "description": "Quantize a floating-point array into int8_t format.",
              "examples": [],
              "id": "arm_quantize_f32_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_quantize_f32_s8",
              "params": [
                {
                  "description": "Pointer to the input float array.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float *"
                },
                {
                  "description": "Pointer to the output int8_t array.",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements in the arrays.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                },
                {
                  "description": "Zero point (offset) to apply during quantization.",
                  "direction": "in",
                  "name": "zero_point",
                  "type": "int32_t"
                },
                {
                  "description": "Scale factor to apply during quantization. Must be a positive finite number. A scale that is zero, negative, NaN, Inf, or small enough that its reciprocal overflows is unsupported, and the result is then unspecified: the scalar and Helium legs are not guaranteed to agree for such a scale. Denormal inputs are likewise unspecified: the Helium leg flushes them to zero, so the two legs are not guaranteed to agree for a denormal input at any scale. Screen denormals if you need them to match.",
                  "direction": "in",
                  "name": "scale",
                  "type": "float"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when `zero_point` lies outside the int8_t range. Values round half away from zero and saturate to the int8_t range after the zero point is applied; NaN maps to `zero_point`."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_quantize_f32_s8(\n    const float *input,\n    int8_t *output,\n    int32_t size,\n    int32_t zero_point,\n    float scale\n)",
              "source": {
                "line": 7356,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7356"
              },
              "summary": "Quantize a floating-point array into int8t format."
            },
            {
              "description": "Quantize a floating-point array into int16_t format.",
              "examples": [],
              "id": "arm_quantize_f32_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_quantize_f32_s16",
              "params": [
                {
                  "description": "Pointer to the input float array.",
                  "direction": "in",
                  "name": "input",
                  "type": "const float *"
                },
                {
                  "description": "Pointer to the output int16_t array.",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements in the arrays.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                },
                {
                  "description": "Zero point (offset) to apply during quantization.",
                  "direction": "in",
                  "name": "zero_point",
                  "type": "int32_t"
                },
                {
                  "description": "Scale factor to apply during quantization. Must be a positive finite number. A scale that is zero, negative, NaN, Inf, or small enough that its reciprocal overflows is unsupported, and the result is then unspecified: the scalar and Helium legs are not guaranteed to agree for such a scale. Denormal inputs are likewise unspecified: the Helium leg flushes them to zero, so the two legs are not guaranteed to agree for a denormal input at any scale. Screen denormals if you need them to match.",
                  "direction": "in",
                  "name": "scale",
                  "type": "float"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "ARM_CMSIS_NN_SUCCESS, or ARM_CMSIS_NN_ARG_ERROR when `zero_point` lies outside the int16_t range. Values round half away from zero and saturate to the int16_t range after the zero point is applied; NaN maps to `zero_point`."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_quantize_f32_s16(\n    const float *input,\n    int16_t *output,\n    int32_t size,\n    int32_t zero_point,\n    float scale\n)",
              "source": {
                "line": 7376,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7376"
              },
              "summary": "Quantize a floating-point array into int16t format."
            },
            {
              "description": "Requantize an int8_t array to another int8_t range with a different scale.",
              "examples": [],
              "id": "arm_requantize_s8_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_requantize_s8_s8",
              "params": [
                {
                  "description": "Pointer to the input int8_t array.",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the output int8_t array.",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "Number of elements in the arrays.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                },
                {
                  "description": "Multiplier used for the scaling operation.",
                  "direction": "in",
                  "name": "effective_scale_multiplier",
                  "type": "int32_t"
                },
                {
                  "description": "Right or left shift (depending on sign) applied after the multiplier.",
                  "direction": "in",
                  "name": "effective_scale_shift",
                  "type": "int32_t"
                },
                {
                  "description": "Zero point of the input data.",
                  "direction": "in",
                  "name": "input_zeropoint",
                  "type": "int32_t"
                },
                {
                  "description": "Zero point of the output data.",
                  "direction": "in",
                  "name": "output_zeropoint",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_requantize_s8_s8(\n    const int8_t *input,\n    int8_t *output,\n    int32_t size,\n    int32_t effective_scale_multiplier,\n    int32_t effective_scale_shift,\n    int32_t input_zeropoint,\n    int32_t output_zeropoint\n)",
              "source": {
                "line": 7390,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7390"
              },
              "summary": "Requantize an int8t array to another int8t range with a different scale."
            },
            {
              "description": "Requantize an int16_t array to another int16_t range with a different scale.",
              "examples": [],
              "id": "arm_requantize_s16_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_requantize_s16_s16",
              "params": [
                {
                  "description": "Pointer to the input int16_t array.",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to the output int16_t array.",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of elements in the arrays.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                },
                {
                  "description": "Multiplier used for the scaling operation.",
                  "direction": "in",
                  "name": "effective_scale_multiplier",
                  "type": "int32_t"
                },
                {
                  "description": "Right or left shift (depending on sign) applied after the multiplier.",
                  "direction": "in",
                  "name": "effective_scale_shift",
                  "type": "int32_t"
                },
                {
                  "description": "Zero point of the input data.",
                  "direction": "in",
                  "name": "input_zeropoint",
                  "type": "int32_t"
                },
                {
                  "description": "Zero point of the output data.",
                  "direction": "in",
                  "name": "output_zeropoint",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_requantize_s16_s16(\n    const int16_t *input,\n    int16_t *output,\n    int32_t size,\n    int32_t effective_scale_multiplier,\n    int32_t effective_scale_shift,\n    int32_t input_zeropoint,\n    int32_t output_zeropoint\n)",
              "source": {
                "line": 7410,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7410"
              },
              "summary": "Requantize an int16t array to another int16t range with a different scale."
            },
            {
              "description": "Dequantize an int8_t array back to floating-point format.",
              "examples": [],
              "id": "arm_dequantize_s8_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_dequantize_s8_f32",
              "params": [
                {
                  "description": "Pointer to the input int8_t array.",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the output float array.",
                  "direction": "out",
                  "name": "output",
                  "type": "float *"
                },
                {
                  "description": "Number of elements in the arrays.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                },
                {
                  "description": "Zero point (offset) that was used during quantization.",
                  "direction": "in",
                  "name": "zero_point",
                  "type": "int32_t"
                },
                {
                  "description": "Scale factor that was used during quantization.",
                  "direction": "in",
                  "name": "scale",
                  "type": "float"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_dequantize_s8_f32(\n    const int8_t *input,\n    float *output,\n    int32_t size,\n    int32_t zero_point,\n    float scale\n)",
              "source": {
                "line": 7429,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7429"
              },
              "summary": "Dequantize an int8t array back to floating-point format."
            },
            {
              "description": "Dequantize an int16_t array back to floating-point format.",
              "examples": [],
              "id": "arm_dequantize_s16_f32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_dequantize_s16_f32",
              "params": [
                {
                  "description": "Pointer to the input int16_t array.",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to the output float array.",
                  "direction": "out",
                  "name": "output",
                  "type": "float *"
                },
                {
                  "description": "Number of elements in the arrays.",
                  "direction": "in",
                  "name": "size",
                  "type": "int32_t"
                },
                {
                  "description": "Zero point (offset) that was used during quantization.",
                  "direction": "in",
                  "name": "zero_point",
                  "type": "int32_t"
                },
                {
                  "description": "Scale factor that was used during quantization.",
                  "direction": "in",
                  "name": "scale",
                  "type": "float"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_dequantize_s16_f32(\n    const int16_t *input,\n    float *output,\n    int32_t size,\n    int32_t zero_point,\n    float scale\n)",
              "source": {
                "line": 7442,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7442"
              },
              "summary": "Dequantize an int16t array back to floating-point format."
            },
            {
              "description": "Strided slice function for int8 data.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_strided_slice_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_strided_slice_s8",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Begin dimensions for slicing",
                  "direction": "in",
                  "name": "begin_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Stride dimensions for slicing",
                  "direction": "in",
                  "name": "stride_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *const"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_strided_slice_s8(\n    const int8_t *input_data,\n    int8_t *output_data,\n    const cmsis_nn_dims *const input_dims,\n    const cmsis_nn_dims *const begin_dims,\n    const cmsis_nn_dims *const stride_dims,\n    const cmsis_nn_dims *const output_dims\n)",
              "source": {
                "line": 7465,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7465"
              },
              "summary": "Strided slice function for int8 data."
            },
            {
              "description": "Strided slice function for int16 data.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_strided_slice_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_strided_slice_s16",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Begin dimensions for slicing",
                  "direction": "in",
                  "name": "begin_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Stride dimensions for slicing",
                  "direction": "in",
                  "name": "stride_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *const"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_strided_slice_s16(\n    const int16_t *input_data,\n    int16_t *output_data,\n    const cmsis_nn_dims *const input_dims,\n    const cmsis_nn_dims *const begin_dims,\n    const cmsis_nn_dims *const stride_dims,\n    const cmsis_nn_dims *const output_dims\n)",
              "source": {
                "line": 7488,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7488"
              },
              "summary": "Strided slice function for int16 data."
            },
            {
              "description": "Strided slice function for int32 data.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_strided_slice_s32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_strided_slice_s32",
              "params": [
                {
                  "description": "Pointer to input tensor",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to output tensor",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int32_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Begin dimensions for slicing",
                  "direction": "in",
                  "name": "begin_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Stride dimensions for slicing",
                  "direction": "in",
                  "name": "stride_dims",
                  "type": "const cmsis_nn_dims *const"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *const"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_strided_slice_s32(\n    const int32_t *input_data,\n    int32_t *output_data,\n    const cmsis_nn_dims *const input_dims,\n    const cmsis_nn_dims *const begin_dims,\n    const cmsis_nn_dims *const stride_dims,\n    const cmsis_nn_dims *const output_dims\n)",
              "source": {
                "line": 7511,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7511"
              },
              "summary": "Strided slice function for int32 data."
            },
            {
              "description": "Gather elements along an axis for int8 tensors.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_gather_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gather_s8",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to indices tensor data (int32)",
                  "direction": "in",
                  "name": "indices_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Indices tensor dimensions",
                  "direction": "in",
                  "name": "indices_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to gather parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gather_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gather_s8(\n    const int8_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t *indices_data,\n    const cmsis_nn_dims *indices_dims,\n    const cmsis_nn_gather_params *params,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 7540,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7540"
              },
              "summary": "Gather elements along an axis for int8 tensors."
            },
            {
              "description": "Gather elements along an axis for int16 tensors.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_gather_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gather_s16",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input tensor dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to indices tensor data (int32)",
                  "direction": "in",
                  "name": "indices_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Indices tensor dimensions",
                  "direction": "in",
                  "name": "indices_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to gather parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gather_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gather_s16(\n    const int16_t *input_data,\n    const cmsis_nn_dims *input_dims,\n    const int32_t *indices_data,\n    const cmsis_nn_dims *indices_dims,\n    const cmsis_nn_gather_params *params,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 7565,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7565"
              },
              "summary": "Gather elements along an axis for int16 tensors."
            },
            {
              "description": "Gather_nd slices for int8 tensors.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_gather_nd_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gather_nd_s8",
              "params": [
                {
                  "description": "Pointer to params tensor data",
                  "direction": "in",
                  "name": "params_data",
                  "type": "const int8_t *"
                },
                {
                  "description": "Params tensor dimensions",
                  "direction": "in",
                  "name": "params_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to indices tensor data (int32)",
                  "direction": "in",
                  "name": "indices_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Indices tensor dimensions",
                  "direction": "in",
                  "name": "indices_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to gather_nd parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gather_nd_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int8_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gather_nd_s8(\n    const int8_t *params_data,\n    const cmsis_nn_dims *params_dims,\n    const int32_t *indices_data,\n    const cmsis_nn_dims *indices_dims,\n    const cmsis_nn_gather_nd_params *params,\n    int8_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 7590,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7590"
              },
              "summary": "Gathernd slices for int8 tensors."
            },
            {
              "description": "Gather_nd slices for int16 tensors.\n\n1. Supported framework: TensorFlow Lite Micro",
              "examples": [],
              "id": "arm_gather_nd_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_gather_nd_s16",
              "params": [
                {
                  "description": "Pointer to params tensor data",
                  "direction": "in",
                  "name": "params_data",
                  "type": "const int16_t *"
                },
                {
                  "description": "Params tensor dimensions",
                  "direction": "in",
                  "name": "params_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to indices tensor data (int32)",
                  "direction": "in",
                  "name": "indices_data",
                  "type": "const int32_t *"
                },
                {
                  "description": "Indices tensor dimensions",
                  "direction": "in",
                  "name": "indices_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Pointer to gather_nd parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_gather_nd_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output_data",
                  "type": "int16_t *"
                },
                {
                  "description": "Output tensor dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_gather_nd_s16(\n    const int16_t *params_data,\n    const cmsis_nn_dims *params_dims,\n    const int32_t *indices_data,\n    const cmsis_nn_dims *indices_dims,\n    const cmsis_nn_gather_nd_params *params,\n    int16_t *output_data,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 7615,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7615"
              },
              "summary": "Gathernd slices for int16 tensors."
            },
            {
              "description": "Tile an int8 tensor along each dimension.\n\n1. Supported framework: TensorFlow Lite Micro\n2. Maximum rank: 8",
              "examples": [],
              "id": "arm_tile_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_tile_s8",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to tile parameters (rank, input_shape, multiples)",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_tile_params *"
                },
                {
                  "description": "Pointer to output tensor data (pre-allocated by caller)",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_tile_s8(const int8_t *input, const cmsis_nn_tile_params *params, int8_t *output)",
              "source": {
                "line": 7642,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7642"
              },
              "summary": "Tile an int8 tensor along each dimension."
            },
            {
              "description": "Tile an int16 tensor along each dimension.",
              "examples": [],
              "id": "arm_tile_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_tile_s16",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to tile parameters (rank, input_shape, multiples)",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_tile_params *"
                },
                {
                  "description": "Pointer to output tensor data (pre-allocated by caller)",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_tile_s16(const int16_t *input, const cmsis_nn_tile_params *params, int16_t *output)",
              "source": {
                "line": 7654,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7654"
              },
              "summary": "Tile an int16 tensor along each dimension."
            },
            {
              "description": "Broadcast an int8 tensor to a target shape.\n\n1. Input dimensions must be 1 or match the output dimension for each axis.\n2. Maximum rank: 8",
              "examples": [],
              "id": "arm_broadcast_to_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_broadcast_to_s8",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to broadcast parameters (rank, input/output shapes)",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_broadcast_to_params *"
                },
                {
                  "description": "Pointer to output tensor data (pre-allocated by caller)",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_broadcast_to_s8(const int8_t *input, const cmsis_nn_broadcast_to_params *params, int8_t *output)",
              "source": {
                "line": 7676,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7676"
              },
              "summary": "Broadcast an int8 tensor to a target shape."
            },
            {
              "description": "Broadcast an int16 tensor to a target shape.",
              "examples": [],
              "id": "arm_broadcast_to_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_broadcast_to_s16",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to broadcast parameters (rank, input/output shapes)",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_broadcast_to_params *"
                },
                {
                  "description": "Pointer to output tensor data (pre-allocated by caller)",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_broadcast_to_s16(\n    const int16_t *input,\n    const cmsis_nn_broadcast_to_params *params,\n    int16_t *output\n)",
              "source": {
                "line": 7689,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7689"
              },
              "summary": "Broadcast an int16 tensor to a target shape."
            },
            {
              "description": "Scatter updates into a zero-initialized output tensor for int8.",
              "examples": [],
              "id": "arm_scatter_nd_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_scatter_nd_s8",
              "params": [
                {
                  "description": "Pointer to indices data (int32, shape [num_updates, index_depth])",
                  "direction": "in",
                  "name": "indices",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to updates data",
                  "direction": "in",
                  "name": "updates",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to scatter_nd parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_scatter_nd_params *"
                },
                {
                  "description": "Pointer to output tensor data (pre-allocated by caller)",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_scatter_nd_s8(\n    const int32_t *indices,\n    const int8_t *updates,\n    const cmsis_nn_scatter_nd_params *params,\n    int8_t *output\n)",
              "source": {
                "line": 7707,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7707"
              },
              "summary": "Scatter updates into a zero-initialized output tensor for int8."
            },
            {
              "description": "Scatter updates into a zero-initialized output tensor for int16.",
              "examples": [],
              "id": "arm_scatter_nd_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_scatter_nd_s16",
              "params": [
                {
                  "description": "Pointer to indices data (int32, shape [num_updates, index_depth])",
                  "direction": "in",
                  "name": "indices",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to updates data",
                  "direction": "in",
                  "name": "updates",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to scatter_nd parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_scatter_nd_params *"
                },
                {
                  "description": "Pointer to output tensor data (pre-allocated by caller)",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_scatter_nd_s16(\n    const int32_t *indices,\n    const int16_t *updates,\n    const cmsis_nn_scatter_nd_params *params,\n    int16_t *output\n)",
              "source": {
                "line": 7723,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7723"
              },
              "summary": "Scatter updates into a zero-initialized output tensor for int16."
            },
            {
              "description": "Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode).",
              "examples": [],
              "id": "arm_mirror_pad_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_mirror_pad_s8",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to mirror_pad parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_mirror_pad_params *"
                },
                {
                  "description": "Pointer to output tensor data (pre-allocated by caller)",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_mirror_pad_s8(const int8_t *input, const cmsis_nn_mirror_pad_params *params, int8_t *output)",
              "source": {
                "line": 7743,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7743"
              },
              "summary": "Mirror-pad an int8 tensor (REFLECT or SYMMETRIC mode)."
            },
            {
              "description": "Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode).",
              "examples": [],
              "id": "arm_mirror_pad_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_mirror_pad_s16",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to mirror_pad parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_mirror_pad_params *"
                },
                {
                  "description": "Pointer to output tensor data (pre-allocated by caller)",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_mirror_pad_s16(const int16_t *input, const cmsis_nn_mirror_pad_params *params, int16_t *output)",
              "source": {
                "line": 7755,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7755"
              },
              "summary": "Mirror-pad an int16 tensor (REFLECT or SYMMETRIC mode)."
            },
            {
              "description": "WHERE operator: return coordinates of non-zero elements in condition.",
              "examples": [],
              "id": "arm_where_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_where_s8",
              "params": [
                {
                  "description": "Pointer to condition tensor data (int8, non-zero = true)",
                  "direction": "in",
                  "name": "condition",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to where parameters (rank, shape)",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_where_params *"
                },
                {
                  "description": "Pointer to output coordinates (int64, shape [max_true, rank])",
                  "direction": "out",
                  "name": "output",
                  "type": "int64_t *"
                },
                {
                  "description": "Number of true elements found",
                  "direction": "out",
                  "name": "num_true",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_where_s8(\n    const int8_t *condition,\n    const cmsis_nn_where_params *params,\n    int64_t *output,\n    int32_t *num_true\n)",
              "source": {
                "line": 7774,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7774"
              },
              "summary": "WHERE operator: return coordinates of non-zero elements in condition."
            },
            {
              "description": "WHERE operator: return coordinates of non-zero elements in condition (int16).",
              "examples": [],
              "id": "arm_where_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_where_s16",
              "params": [
                {
                  "description": "Pointer to condition tensor data (int16, non-zero = true)",
                  "direction": "in",
                  "name": "condition",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to where parameters (rank, shape)",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_where_params *"
                },
                {
                  "description": "Pointer to output coordinates (int64, shape [max_true, rank])",
                  "direction": "out",
                  "name": "output",
                  "type": "int64_t *"
                },
                {
                  "description": "Number of true elements found",
                  "direction": "out",
                  "name": "num_true",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_where_s16(\n    const int16_t *condition,\n    const cmsis_nn_where_params *params,\n    int64_t *output,\n    int32_t *num_true\n)",
              "source": {
                "line": 7788,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7788"
              },
              "summary": "WHERE operator: return coordinates of non-zero elements in condition (int16)."
            },
            {
              "description": "SELECT_V2 with broadcast for int8 tensors.",
              "examples": [],
              "id": "arm_select_v2_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_select_v2_s8",
              "params": [
                {
                  "description": "Pointer to condition tensor data (bool)",
                  "direction": "in",
                  "name": "condition",
                  "type": "const bool *"
                },
                {
                  "description": "Pointer to x tensor data (selected when condition is true)",
                  "direction": "in",
                  "name": "x",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to y tensor data (selected when condition is false)",
                  "direction": "in",
                  "name": "y",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to select_v2 parameters (broadcast strides)",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_select_v2_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_select_v2_s8(\n    const bool *condition,\n    const int8_t *x,\n    const int8_t *y,\n    const cmsis_nn_select_v2_params *params,\n    int8_t *output\n)",
              "source": {
                "line": 7802,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7802"
              },
              "summary": "SELECTV2 with broadcast for int8 tensors."
            },
            {
              "description": "SELECT_V2 with broadcast for int16 tensors.",
              "examples": [],
              "id": "arm_select_v2_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_select_v2_s16",
              "params": [
                {
                  "description": "Pointer to condition tensor data (bool)",
                  "direction": "in",
                  "name": "condition",
                  "type": "const bool *"
                },
                {
                  "description": "Pointer to x tensor data",
                  "direction": "in",
                  "name": "x",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to y tensor data",
                  "direction": "in",
                  "name": "y",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to select_v2 parameters (broadcast strides)",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_select_v2_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_select_v2_s16(\n    const bool *condition,\n    const int16_t *x,\n    const int16_t *y,\n    const cmsis_nn_select_v2_params *params,\n    int16_t *output\n)",
              "source": {
                "line": 7820,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7820"
              },
              "summary": "SELECTV2 with broadcast for int16 tensors."
            },
            {
              "description": "Reverse variable-length sequences along a dimension for int8.",
              "examples": [],
              "id": "arm_reverse_sequence_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reverse_sequence_s8",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to per-batch sequence lengths (int32)",
                  "direction": "in",
                  "name": "seq_lengths",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to reverse_sequence parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_reverse_sequence_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reverse_sequence_s8(\n    const int8_t *input,\n    const int32_t *seq_lengths,\n    const cmsis_nn_reverse_sequence_params *params,\n    int8_t *output\n)",
              "source": {
                "line": 7842,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7842"
              },
              "summary": "Reverse variable-length sequences along a dimension for int8."
            },
            {
              "description": "Reverse variable-length sequences along a dimension for int16.",
              "examples": [],
              "id": "arm_reverse_sequence_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reverse_sequence_s16",
              "params": [
                {
                  "description": "Pointer to input tensor data",
                  "direction": "in",
                  "name": "input",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to per-batch sequence lengths (int32)",
                  "direction": "in",
                  "name": "seq_lengths",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to reverse_sequence parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_reverse_sequence_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_reverse_sequence_s16(\n    const int16_t *input,\n    const int32_t *seq_lengths,\n    const cmsis_nn_reverse_sequence_params *params,\n    int16_t *output\n)",
              "source": {
                "line": 7858,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7858"
              },
              "summary": "Reverse variable-length sequences along a dimension for int16."
            },
            {
              "description": "Update a slice of an int8 operand tensor at runtime-determined indices.",
              "examples": [],
              "id": "arm_dynamic_update_slice_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_dynamic_update_slice_s8",
              "params": [
                {
                  "description": "Pointer to operand tensor data (copied to output first)",
                  "direction": "in",
                  "name": "operand",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to update tensor data",
                  "direction": "in",
                  "name": "update",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to start index per dimension (int32, length = rank)",
                  "direction": "in",
                  "name": "start_indices",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to dynamic_update_slice parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_dynamic_update_slice_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_dynamic_update_slice_s8(\n    const int8_t *operand,\n    const int8_t *update,\n    const int32_t *start_indices,\n    const cmsis_nn_dynamic_update_slice_params *params,\n    int8_t *output\n)",
              "source": {
                "line": 7880,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7880"
              },
              "summary": "Update a slice of an int8 operand tensor at runtime-determined indices."
            },
            {
              "description": "Update a slice of an int16 operand tensor at runtime-determined indices.",
              "examples": [],
              "id": "arm_dynamic_update_slice_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_dynamic_update_slice_s16",
              "params": [
                {
                  "description": "Pointer to operand tensor data (copied to output first)",
                  "direction": "in",
                  "name": "operand",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to update tensor data",
                  "direction": "in",
                  "name": "update",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to start index per dimension (int32, length = rank)",
                  "direction": "in",
                  "name": "start_indices",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to dynamic_update_slice parameters",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_dynamic_update_slice_params *"
                },
                {
                  "description": "Pointer to output tensor data",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_dynamic_update_slice_s16(\n    const int16_t *operand,\n    const int16_t *update,\n    const int32_t *start_indices,\n    const cmsis_nn_dynamic_update_slice_params *params,\n    int16_t *output\n)",
              "source": {
                "line": 7898,
                "path": "Include/arm_nnfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions.h#L7898"
              },
              "summary": "Update a slice of an int16 operand tensor at runtime-determined indices."
            }
          ]
        },
        {
          "description": "",
          "name": "arm_nnsupportfunctions.h",
          "path": "heliaCORE.arm_nnsupportfunctions",
          "submodules": [],
          "summary": "",
          "symbols": [
            {
              "description": "",
              "examples": [],
              "id": "USE_FAST_DW_CONV_S16_FUNCTION",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "USE_FAST_DW_CONV_S16_FUNCTION",
              "params": [
                {
                  "description": "",
                  "name": "dw_conv_params"
                },
                {
                  "description": "",
                  "name": "filter_dims"
                },
                {
                  "description": "",
                  "name": "input_dims"
                },
                {
                  "description": "",
                  "name": "output_dims"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define USE_FAST_DW_CONV_S16_FUNCTION(dw_conv_params, filter_dims, input_dims, output_dims) (dw_conv_params->ch_mult == 1 && \\ arm_nn_dw_conv_opt_dilation_supported(dw_conv_params, input_dims, filter_dims, output_dims) && \\ filter_dims->w * filter_dims->h < 512)",
              "source": {
                "line": 45,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L45"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "LEFT_SHIFT",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "LEFT_SHIFT",
              "params": [
                {
                  "description": "",
                  "name": "_shift"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define LEFT_SHIFT(_shift) (_shift > 0 ? _shift : 0)",
              "source": {
                "line": 50,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L50"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "RIGHT_SHIFT",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "RIGHT_SHIFT",
              "params": [
                {
                  "description": "",
                  "name": "_shift"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define RIGHT_SHIFT(_shift) (_shift > 0 ? 0 : -_shift)",
              "source": {
                "line": 51,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L51"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "MASK_IF_ZERO",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "MASK_IF_ZERO",
              "params": [
                {
                  "description": "",
                  "name": "x"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define MASK_IF_ZERO(x) (x) == 0 ? ~0 : 0",
              "source": {
                "line": 52,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L52"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "MASK_IF_NON_ZERO",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "MASK_IF_NON_ZERO",
              "params": [
                {
                  "description": "",
                  "name": "x"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define MASK_IF_NON_ZERO(x) (x) != 0 ? ~0 : 0",
              "source": {
                "line": 53,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L53"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "SELECT_USING_MASK",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "SELECT_USING_MASK",
              "params": [
                {
                  "description": "",
                  "name": "mask"
                },
                {
                  "description": "",
                  "name": "a"
                },
                {
                  "description": "",
                  "name": "b"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define SELECT_USING_MASK(mask, a, b) ((mask) & (a)) ^ (~(mask) & (b))",
              "source": {
                "line": 54,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L54"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "ARM_NN_MAX",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_MAX",
              "params": [
                {
                  "description": "",
                  "name": "A"
                },
                {
                  "description": "",
                  "name": "B"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_MAX(A, B) ((A) > (B) ? (A) : (B))",
              "source": {
                "line": 57,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L57"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "ARM_NN_MIN",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_MIN",
              "params": [
                {
                  "description": "",
                  "name": "A"
                },
                {
                  "description": "",
                  "name": "B"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_MIN(A, B) ((A) < (B) ? (A) : (B))",
              "source": {
                "line": 58,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L58"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "ARM_NN_CLAMP",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_CLAMP",
              "params": [
                {
                  "description": "",
                  "name": "x"
                },
                {
                  "description": "",
                  "name": "h"
                },
                {
                  "description": "",
                  "name": "l"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_CLAMP(x, h, l) ARM_NN_MAX(ARM_NN_MIN((x), (h)), (l))",
              "source": {
                "line": 59,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L59"
              },
              "summary": ""
            },
            {
              "description": "Minimum of two scalar f16 values.\n\nWith ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 minNum via VMINNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MIN, an ordered compare, so its NaN handling depends on operand order: a NaN `b` is returned, a NaN `a` is not. Do not rely on NaN suppression on non-MVE builds.",
              "examples": [],
              "id": "arm_nn_min_f16h",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_min_f16h",
              "params": [
                {
                  "description": "First operand",
                  "direction": "in",
                  "name": "a",
                  "type": "_Float16"
                },
                {
                  "description": "Second operand",
                  "direction": "in",
                  "name": "b",
                  "type": "_Float16"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The smaller of `a` and `b`"
                }
              ],
              "signature": "static _Float16 arm_nn_min_f16h(_Float16 a, _Float16 b)",
              "source": {
                "line": 98,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L98"
              },
              "summary": "Minimum of two scalar f16 values."
            },
            {
              "description": "Maximum of two scalar f16 values.\n\nWith ARM_NN_F16_CMOV_WORKAROUND this is IEEE 754 maxNum via VMAXNM.F16: a NaN operand is suppressed and the non-NaN operand wins. The scalar fallback is ARM_NN_MAX, an ordered compare, so its NaN handling depends on operand order: a NaN `b` is returned, a NaN `a` is not. Do not rely on NaN suppression on non-MVE builds.",
              "examples": [],
              "id": "arm_nn_max_f16h",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_max_f16h",
              "params": [
                {
                  "description": "First operand",
                  "direction": "in",
                  "name": "a",
                  "type": "_Float16"
                },
                {
                  "description": "Second operand",
                  "direction": "in",
                  "name": "b",
                  "type": "_Float16"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The larger of `a` and `b`"
                }
              ],
              "signature": "static _Float16 arm_nn_max_f16h(_Float16 a, _Float16 b)",
              "source": {
                "line": 121,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L121"
              },
              "summary": "Maximum of two scalar f16 values."
            },
            {
              "description": "Returns `x` when `x` is NaN, otherwise `y`.\n\nBoth the NaN test and the select are performed on the bit patterns: the test is (bits & 0x7FFF) > 0x7C00 (all-ones exponent, non-zero mantissa), which is integer arithmetic that -ffinite-math-only (implied by the shipped -Ofast) has no license to fold, unlike the former floating-point self-compare `x != x` (#333 / #334); the bit-pattern select neither expands to an HFmode conditional move (PR target/118460) nor quiets/retags the NaN payload. This helper backs the f16 elementwise clamp and, via arm_nn_clamp_scalar_f16 / arm_nn_clamp_propagate_nan_f16h, the other f16 scalar clamp users  arm_svdf_f16, arm_max_pool_f16 / arm_avg_pool_f16, the packed f16 matmul (arm_nn_mat_mult_nt_n_packed_f16), the scalar f16 RELU/RELU6/LEAKY_RELU activation legs, and arm_nn_vector_clamp_f16's scalar leg (conv/depthwise/transpose-conv f16, the 3x3 depthwise, and arm_nn_maxpool1d_f16)  so wherever that scalar clamp runs, a NaN passes through it at every optimization level on the gated toolchains. That is a guarantee about the clamp, not the whole kernel: which builds run the scalar clamp, and whether a NaN survives the rest of the kernel to reach it, is per kernel  several of these callers clamp with vmaxnmq/vminnmq on MVE builds (a NaN resolves to a bound there), and arm_max_pool_f16's max reduction drops a NaN before the clamp. The kernels with a NaN\n\n:::note\n(svdf, max/avg pool, packed matmul, activation) state their exact scope there; the arm_nn_vector_clamp_f16 family is covered by a test assertion in the transpose-conv f16 suite rather than per-kernel notes. The cortex-m55 MVE RELU/RELU6 f16 legs reach the same guarantee by the vector form of this idiom rather than by calling this helper: they restore NaN lanes with arm_nn_max_propagate_nan_mve_f16 / arm_nn_clamp_propagate_nan_mve_f16 (#382). The same idiom (bit-classified select) appears in arm_prelu_f16, which does not call this helper.\n\n:::",
              "examples": [],
              "id": "arm_nn_propagate_nan_f16h",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_propagate_nan_f16h",
              "params": [
                {
                  "description": "Value whose NaN-ness selects the result. Returned unchanged when it is NaN.",
                  "direction": "in",
                  "name": "x",
                  "type": "_Float16"
                },
                {
                  "description": "Value returned when `x` is not NaN",
                  "direction": "in",
                  "name": "y",
                  "type": "_Float16"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`x` if `x` is NaN, otherwise `y`"
                }
              ],
              "signature": "static _Float16 arm_nn_propagate_nan_f16h(_Float16 x, _Float16 y)",
              "source": {
                "line": 167,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L167"
              },
              "summary": "Returns x when x is NaN, otherwise y."
            },
            {
              "description": "Drop-in equivalent of `ARM_NN_CLAMP(x, h, l)` for scalar _Float16 operands.\n\nIncludes the macro's NaN behaviour: `ARM_NN_MIN(NaN, h)` is h, so a NaN input resolves to the high bound, exactly as the macro does. Use arm_nn_clamp_propagate_nan_f16h() where TFLite NaN propagation is required.",
              "examples": [],
              "id": "arm_nn_clamp_f16h",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_clamp_f16h",
              "params": [
                {
                  "description": "Value to clamp",
                  "direction": "in",
                  "name": "x",
                  "type": "_Float16"
                },
                {
                  "description": "Upper bound",
                  "direction": "in",
                  "name": "h",
                  "type": "_Float16"
                },
                {
                  "description": "Lower bound",
                  "direction": "in",
                  "name": "l",
                  "type": "_Float16"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`x` clamped to [`l`, `h`]"
                }
              ],
              "signature": "static _Float16 arm_nn_clamp_f16h(_Float16 x, _Float16 h, _Float16 l)",
              "source": {
                "line": 193,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L193"
              },
              "summary": "Drop-in equivalent of ARMNNCLAMP(x, h, l) for scalar Float16 operands."
            },
            {
              "description": "Scalar f16 clamp with TFLite NaN semantics: NaN passes through unchanged.\n\nMirrors the MVE idiom in arm_nn_clamp_propagate_nan_mve_f16() (lower bound first, then upper bound, then restore NaN lanes). The NaN restore in arm_nn_propagate_nan_f16h() tests the integer bit pattern, so it holds at every optimization level including the shipped -Ofast; see #333 / #334. Bounds are assumed ordered (l <= h); inverted bounds are unspecified.",
              "examples": [],
              "id": "arm_nn_clamp_propagate_nan_f16h",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_clamp_propagate_nan_f16h",
              "params": [
                {
                  "description": "Value to clamp",
                  "direction": "in",
                  "name": "x",
                  "type": "_Float16"
                },
                {
                  "description": "Lower bound",
                  "direction": "in",
                  "name": "l",
                  "type": "_Float16"
                },
                {
                  "description": "Upper bound",
                  "direction": "in",
                  "name": "h",
                  "type": "_Float16"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`x` clamped to [`l`, `h`], or `x` itself when it is NaN"
                }
              ],
              "signature": "static _Float16 arm_nn_clamp_propagate_nan_f16h(_Float16 x, _Float16 l, _Float16 h)",
              "source": {
                "line": 212,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L212"
              },
              "summary": "Scalar f16 clamp with TFLite NaN semantics: NaN passes through unchanged."
            },
            {
              "description": "Absolute value of a scalar f16 value.",
              "examples": [],
              "id": "arm_nn_abs_f16h",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_abs_f16h",
              "params": [
                {
                  "description": "Input value",
                  "direction": "in",
                  "name": "x",
                  "type": "_Float16"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "|`x`|"
                }
              ],
              "signature": "static _Float16 arm_nn_abs_f16h(_Float16 x)",
              "source": {
                "line": 224,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L224"
              },
              "summary": "Absolute value of a scalar f16 value."
            },
            {
              "description": "",
              "examples": [],
              "id": "ARM_NN_ROUND_UP",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_ROUND_UP",
              "params": [
                {
                  "description": "",
                  "name": "x"
                },
                {
                  "description": "",
                  "name": "multiple"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_ROUND_UP(x, multiple) ((((x) + (multiple) - 1) / (multiple)) * (multiple))",
              "source": {
                "line": 236,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L236"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "REDUCE_MULTIPLIER",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "REDUCE_MULTIPLIER",
              "params": [
                {
                  "description": "",
                  "name": "_mult"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define REDUCE_MULTIPLIER(_mult) ((_mult < 0x7FFF0000) ? ((_mult + (1 << 15)) >> 16) : 0x7FFF)",
              "source": {
                "line": 237,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L237"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "CH_IN_BLOCK_MVE",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "CH_IN_BLOCK_MVE",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define CH_IN_BLOCK_MVE (124)",
              "source": {
                "line": 245,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L245"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "S4_CH_IN_BLOCK_MVE",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "S4_CH_IN_BLOCK_MVE",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define S4_CH_IN_BLOCK_MVE (124)",
              "source": {
                "line": 250,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L250"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "MAX_COL_COUNT",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "MAX_COL_COUNT",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define MAX_COL_COUNT (512)",
              "source": {
                "line": 254,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L254"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "REVERSE_TCOL_EFFICIENT_THRESHOLD",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "REVERSE_TCOL_EFFICIENT_THRESHOLD",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define REVERSE_TCOL_EFFICIENT_THRESHOLD (16)",
              "source": {
                "line": 258,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L258"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD (1)",
              "source": {
                "line": 266,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L266"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "OPTIONAL_RESTRICT_KEYWORD",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "OPTIONAL_RESTRICT_KEYWORD",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define OPTIONAL_RESTRICT_KEYWORD",
              "source": {
                "line": 272,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L272"
              },
              "summary": ""
            },
            {
              "description": "Fold one dimension into a running buffer-size product, reporting overflow as -1.\n\nBuffer-size queries return an int32_t byte count, so the product of the dimensions they multiply has to be rejected as soon as it cannot fit. Folding one factor at a time keeps the accumulator bounded: an accumulator already known to be <= INT32_MAX times a factor <= INT32_MAX cannot exceed about 2^62, so the int64_t accumulator itself never wraps. Chaining raw (int64_t) casts across three or more int32_t dims does not have that property - 65536 * 65536 * 65536 * 65536 is exactly 2^64 and folds back to 0, which would sail through a trailing \"> INT32_MAX\" test.\n\n:::note\nThis is the -1 sentinel family, used by the s8/s16 integer buffer-size queries, by the eight SVDF staging queries (arm_svdf_{s8,state_s16_s8,f32,f16}_{input,output}_ctx_get_buffer_size) and by the s8/s16 LSTM temp-buffer queries and the GRU temp queries (arm_lstm_unidirectional_{s8,s16}_temp{1,2}_get_buffer_size, arm_gru_unidirectional_{f32,f16}_temp1_get_buffer_size). The four f32/f16 LSTM temp queries have no dimensions to fold (the buffers are unused) and answer -1 only for NULL params, 0 otherwise. It is not interchangeable with the arm_nn_checked_size_mul() / arm_nn_size_to_i32_or_zero() helpers in Source/NNSupportFunctions (shared header for the float sizers), which most f32 and f16 buffer-size queries use and which report an out-of-range size as 0. Mixing the two silently flips a sizer's out-of-range contract from \"must never be used to size a buffer\" to \"you may pass { NULL, 0 }\", so pick the one the surrounding family already uses.\n\n:::\n\n:::note\nThe split is per sizer, not per datatype. The four SVDF f32/f16 staging queries deliberately use this -1 family rather than the 0 one their neighbours use, because their kernels read ctx->size and a size of 0 opts out of the scratch-size check - so a 0-on-overflow answer fed back as { alloc(0), 0 } would disable the very check meant to catch it. Do not infer a sizer's sentinel from its datatype suffix.\n\n:::",
              "examples": [],
              "id": "arm_nn_size_mul",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_size_mul",
              "params": [
                {
                  "description": "Running product, or -1 if an earlier fold already overflowed.",
                  "direction": "in",
                  "name": "acc",
                  "type": "const int64_t"
                },
                {
                  "description": "Next factor to fold in.",
                  "direction": "in",
                  "name": "factor",
                  "type": "const int64_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "acc * factor, or -1 if acc is already -1, factor is negative or out of int32_t range, or the product exceeds INT32_MAX."
                }
              ],
              "signature": "static int64_t arm_nn_size_mul(const int64_t acc, const int64_t factor)",
              "source": {
                "line": 307,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L307"
              },
              "summary": "Fold one dimension into a running buffer-size product, reporting overflow as -1."
            },
            {
              "description": "Add to a running buffer-size product, reporting overflow as -1.\n\nCompanion to arm_nn_size_mul() for the sizers that append a fixed slack term.\n\n:::note\nSame sentinel caveat as arm_nn_size_mul() - see its note for the full split, including why the four SVDF f32/f16 staging queries use this -1 family rather than the 0-returning arm_nn_checked_size_mul() / arm_nn_size_to_i32_or_zero() family that most other float sizers use.\n\n:::",
              "examples": [],
              "id": "arm_nn_size_add",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_size_add",
              "params": [
                {
                  "description": "Running product, or -1 if an earlier step already overflowed.",
                  "direction": "in",
                  "name": "acc",
                  "type": "const int64_t"
                },
                {
                  "description": "Value to add. Must be non-negative.",
                  "direction": "in",
                  "name": "addend",
                  "type": "const int64_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "acc + addend, or -1 if acc is already -1 or the sum exceeds INT32_MAX."
                }
              ],
              "signature": "static int64_t arm_nn_size_add(const int64_t acc, const int64_t addend)",
              "source": {
                "line": 332,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L332"
              },
              "summary": "Add to a running buffer-size product, reporting overflow as -1."
            },
            {
              "description": "definition to pack four 8 bit values.\n\nByte lanes are masked and shifted in uint32_t so a negative value never feeds a signed left shift (UB); masking before the shift keeps the same bits the old shift-then-mask form kept. Bit-identical for every input. Deliberate divergence from upstream ARM-software/CMSIS-NN, which still carries the signed-shift form  do not paste the upstream text back on a sync (issue #357).",
              "examples": [],
              "id": "PACK_S8x4_32x1",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "PACK_S8x4_32x1",
              "params": [
                {
                  "description": "",
                  "name": "v0"
                },
                {
                  "description": "",
                  "name": "v1"
                },
                {
                  "description": "",
                  "name": "v2"
                },
                {
                  "description": "",
                  "name": "v3"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define PACK_S8x4_32x1(v0, v1, v2, v3) ((int32_t)((((uint32_t)(v0)) & 0xFFu) | ((((uint32_t)(v1)) & 0xFFu) << 8) | ((((uint32_t)(v2)) & 0xFFu) << 16) | \\ ((((uint32_t)(v3)) & 0xFFu) << 24)))",
              "source": {
                "line": 358,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L358"
              },
              "summary": "definition to pack four 8 bit values."
            },
            {
              "description": "definition to pack two 16 bit values.\n\nSame treatment: the high half is shifted in uint32_t, not int32_t, so a negative v1 is defined; the low half keeps its mask. Bit-identical for every input. Same deliberate upstream divergence as PACK_S8x4_32x1 above.",
              "examples": [],
              "id": "PACK_Q15x2_32x1",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "PACK_Q15x2_32x1",
              "params": [
                {
                  "description": "",
                  "name": "v0"
                },
                {
                  "description": "",
                  "name": "v1"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define PACK_Q15x2_32x1(v0, v1) ((int32_t)((((uint32_t)(v0)) & 0xFFFFu) | (((uint32_t)(v1)) << 16)))",
              "source": {
                "line": 369,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L369"
              },
              "summary": "definition to pack two 16 bit values."
            },
            {
              "description": "Map an output index to the nearest input index for resize.\n\nThis helper follows the TensorFlow Lite nearest-neighbor resize mapping rules.",
              "examples": [],
              "id": "GetNearestNeighbor",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "GetNearestNeighbor",
              "params": [
                {
                  "description": "Output index (x or y).",
                  "direction": "in",
                  "name": "input_value",
                  "type": "const int"
                },
                {
                  "description": "Input size along the same axis.",
                  "direction": "in",
                  "name": "input_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Precomputed scaling factor for the axis.",
                  "direction": "in",
                  "name": "scale",
                  "type": "const float"
                },
                {
                  "description": "Precomputed offset for the axis.",
                  "direction": "in",
                  "name": "offset",
                  "type": "const float"
                },
                {
                  "description": "If true, use align-corners scaling.",
                  "direction": "in",
                  "name": "align_corners",
                  "type": "const bool"
                },
                {
                  "description": "If true, use half-pixel center offset.",
                  "direction": "in",
                  "name": "half_pixel_centers",
                  "type": "const bool"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Nearest input index for the given output index."
                }
              ],
              "signature": "static int32_t GetNearestNeighbor(\n    const int input_value,\n    const int32_t input_size,\n    const float scale,\n    const float offset,\n    const bool align_corners,\n    const bool half_pixel_centers\n)",
              "source": {
                "line": 391,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L391"
              },
              "summary": "Map an output index to the nearest input index for resize."
            },
            {
              "description": "Check if convolution parameters correspond to a 1x1 convolution.",
              "examples": [],
              "id": "arm_nn_is_convolve_1x1",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_is_convolve_1x1",
              "params": [
                {
                  "description": "Convolution parameters",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "true if parameters describe a 1x1 convolution, false otherwise."
                }
              ],
              "signature": "static bool arm_nn_is_convolve_1x1(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims\n)",
              "source": {
                "line": 416,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L416"
              },
              "summary": "Check if convolution parameters correspond to a 1x1 convolution."
            },
            {
              "description": "Check if a 1x1 convolution qualifies for the fast (unit stride) path.\n\n:::note\nDoes not validate that the kernel is 1x1. Call arm_nn_is_convolve_1x1() first.\n\n:::",
              "examples": [],
              "id": "arm_nn_is_convolve_1x1_fast",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_is_convolve_1x1_fast",
              "params": [
                {
                  "description": "Convolution parameters",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "true if stride is 1x1, false otherwise."
                }
              ],
              "signature": "static bool arm_nn_is_convolve_1x1_fast(const cmsis_nn_conv_params *conv_params)",
              "source": {
                "line": 432,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L432"
              },
              "summary": "Check if a 1x1 convolution qualifies for the fast (unit stride) path."
            },
            {
              "description": "Check if convolution parameters correspond to a 1xN convolution.",
              "examples": [],
              "id": "arm_nn_is_convolve_1_x_n",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_is_convolve_1_x_n",
              "params": [
                {
                  "description": "Convolution parameters",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "true if parameters describe a 1xN convolution, false otherwise."
                }
              ],
              "signature": "static bool arm_nn_is_convolve_1_x_n(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims\n)",
              "source": {
                "line": 444,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L444"
              },
              "summary": "Check if convolution parameters correspond to a 1xN convolution."
            },
            {
              "description": "Check that `arm_convolve_1_x_n_s4()` handles the horizontal padding of a 1xN convolution.\n\nThe kernel places pad.w columns on the left and pad.w + (total_pad % 2) on the right, where total_pad = (output W - 1) * stride.w + filter W - input W, and needs the output columns that read padding to fit in output W. Its padded-column code also assumes that each such column reads at least one input column and that the filter is no wider than the input; otherwise it forms input and filter addresses outside the tensors. On MVE builds another pad placement, or too many padded columns, returns ARM_CMSIS_NN_FAILURE. A VALID layer whose stride leaves trailing input unused (negative total_pad) is therefore rejected, and the wrapper routes it to another convolution. The kernel also computes a single output row, so vertical padding or an output height other than 1 is rejected. A non-positive stride.w is left to the kernel's argument checks.",
              "examples": [],
              "id": "arm_nn_convolve_1_x_n_padding_supported",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_convolve_1_x_n_padding_supported",
              "params": [
                {
                  "description": "Convolution parameters",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "true when `arm_convolve_1_x_n_s4()` handles the padding, false otherwise."
                }
              ],
              "signature": "static bool arm_nn_convolve_1_x_n_padding_supported(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 473,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L473"
              },
              "summary": "Check that armconvolve1xns4() handles the horizontal padding of a 1xN convolution."
            },
            {
              "description": "Check that `arm_convolve_1_x_n_s8()` accepts the padding and output shape of a 1xN convolution.\n\nThe kernel computes a single output row for any pad.w >= 0 and any output width, including an odd total padding, a filter wider than the input and a VALID layer whose stride leaves trailing input unused. It rejects vertical padding, an output height other than 1, a negative pad.w and an empty filter; the wrapper routes those layers to another convolution. A non-positive stride.w is left to the kernel's argument checks.",
              "examples": [],
              "id": "arm_nn_convolve_1_x_n_s8_padding_supported",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_convolve_1_x_n_s8_padding_supported",
              "params": [
                {
                  "description": "Convolution parameters",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "true when `arm_convolve_1_x_n_s8()` computes the layer, false otherwise."
                }
              ],
              "signature": "static bool arm_nn_convolve_1_x_n_s8_padding_supported(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 517,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L517"
              },
              "summary": "Check that armconvolve1xns8() accepts the padding and output shape of a 1xN convolution."
            },
            {
              "description": "Count the output columns of a 1xN convolution whose window reads padding.\n\nOutput column j reads input columns j * stride.w - pad.w to j * stride.w - pad.w + filter W - 1. The leading columns whose window starts before the input are left-padded; of the others, the trailing columns whose window ends past the input are right-padded. A window can do both only when it is left-padded.",
              "examples": [],
              "id": "arm_nn_convolve_1_x_n_padded_columns",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_convolve_1_x_n_padded_columns",
              "params": [
                {
                  "description": "Convolution parameters. stride.w >= 1 and pad.w >= 0.",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input dimensions. w >= 0.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions. w >= 1.",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output dimensions. w >= 0.",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Number of left-padded output columns, output W at most.",
                  "direction": "out",
                  "name": "left_num",
                  "type": "int64_t *"
                },
                {
                  "description": "Number of right-padded output columns, output W - left_num at most.",
                  "direction": "out",
                  "name": "right_num",
                  "type": "int64_t *"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_nn_convolve_1_x_n_padded_columns(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    int64_t *left_num,\n    int64_t *right_num\n)",
              "source": {
                "line": 539,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L539"
              },
              "summary": "Count the output columns of a 1xN convolution whose window reads padding."
            },
            {
              "description": "Check if the dilation, stride and padding of a depthwise layer allow the `arm_depthwise_conv_s8_opt()` or `arm_depthwise_conv_fast_s16()` route.\n\n:::note\nDoes not check ch_mult, the batch count or the kernel size: `arm_depthwise_conv_wrapper_s8()`, `arm_depthwise_conv_wrapper_s16()` and their buffer-size functions apply their own conditions on those, and all of them take this predicate so that routing and sizing agree.\n\n:::",
              "examples": [],
              "id": "arm_nn_dw_conv_opt_dilation_supported",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_dw_conv_opt_dilation_supported",
              "params": [
                {
                  "description": "Depthwise convolution parameters",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input dimensions",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter dimensions",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output dimensions",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "true for an undilated layer (dilation 1 in both dimensions), or for a 1D layer dilated along the width only: filter, input and output height 1, stride 1 in both dimensions, no vertical padding, dilation.h == 1 and dilation.w >= 1. false otherwise."
                }
              ],
              "signature": "static bool arm_nn_dw_conv_opt_dilation_supported(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 572,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L572"
              },
              "summary": "Check if the dilation, stride and padding of a depthwise layer allow the armdepthwiseconvs8opt() or armdepthwiseconvfasts16() route."
            },
            {
              "description": "Converts the elements from a s8 vector to a s16 vector with an added offset.\n\nOutput elements are ordered. The equation used for the conversion process is:\n\ndst[n] = (int16_t) src[n] + offset; 0 <= n < block_size.",
              "examples": [],
              "id": "arm_q7_to_q15_with_offset",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_q7_to_q15_with_offset",
              "params": [
                {
                  "description": "pointer to the s8 input vector",
                  "direction": "in",
                  "name": "src",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to the s16 output vector",
                  "direction": "out",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "length of the input vector",
                  "direction": "in",
                  "name": "block_size",
                  "type": "int32_t"
                },
                {
                  "description": "s16 offset to be added to each input vector element.",
                  "direction": "in",
                  "name": "offset",
                  "type": "int16_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_q7_to_q15_with_offset(const int8_t *src, int16_t *dst, int32_t block_size, int16_t offset)",
              "source": {
                "line": 649,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L649"
              },
              "summary": "Converts the elements from a s8 vector to a s16 vector with an added offset."
            },
            {
              "description": "Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with MVE extension.\n\nThe dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_s8_opt_get_buffer_size()`. Note also this is a support function, so not recommended to call directly even on Host.\n\n:::\n\n:::note\nThis leg sizes its buffer from a fixed channel block rather than from input_dims->c, but it applies the same dimension check as `arm_depthwise_conv_s8_opt_get_buffer_size()` anyway, returning -1 for a negative input_dims->c or filter dimension, or for a byte count that would not fit in an int32_t. Without that check this entry point - and every s4 depthwise sizer, which route here - answered a negative channel count with a plausible positive size (issue #318).\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt_get_buffer_size_mve",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt_get_buffer_size_mve",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_depthwise_conv_s8_opt_get_buffer_size_mve(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 695,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L695"
              },
              "summary": "Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel."
            },
            {
              "description": "Get the required buffer size for optimized s8 depthwise convolution function with constraint that in_channel equals out_channel. This is for processors with DSP extension.\n\nThe dimensions are checked here, so a negative dimension returns -1 on every build target. The byte count is range-checked inside the selected leg instead, because the Helium and DSP legs use different formulas and the plain-C build needs no buffer at all.\n\n:::note\nIntended for compilation on Host. If compiling for an Arm target, use `arm_depthwise_conv_s8_opt_get_buffer_size()`. Note also this is a support function, so not recommended to call directly even on Host.\n\n:::",
              "examples": [],
              "id": "arm_depthwise_conv_s8_opt_get_buffer_size_dsp",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_depthwise_conv_s8_opt_get_buffer_size_dsp",
              "params": [
                {
                  "description": "Input (activation) tensor dimensions. Format: [1, H, W, C_IN] Batch argument N is not used.",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns required buffer size in bytes, or -1 if any dimension it reads is negative or the required size would not fit in an int32_t"
                }
              ],
              "signature": "int32_t arm_depthwise_conv_s8_opt_get_buffer_size_dsp(const cmsis_nn_dims *input_dims, const cmsis_nn_dims *filter_dims)",
              "source": {
                "line": 710,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L710"
              },
              "summary": "Get the required buffer size for optimized s8 depthwise convolution function with constraint that inchannel equals outchannel."
            },
            {
              "description": "Depthwise conv on an im2col buffer where the input channel equals output channel.\n\nSupported framework: TensorFlow Lite micro.",
              "examples": [],
              "id": "arm_nn_depthwise_conv_s8_core",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv_s8_core",
              "params": [
                {
                  "description": "pointer to row",
                  "direction": "in",
                  "name": "row",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to im2col buffer, always consists of 2 columns.",
                  "direction": "in",
                  "name": "col",
                  "type": "const int16_t *"
                },
                {
                  "description": "number of channels",
                  "direction": "in",
                  "name": "num_ch",
                  "type": "const uint16_t"
                },
                {
                  "description": "pointer to per output channel requantization shift parameter.",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "pointer to per output channel requantization multiplier parameter.",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "output tensor offset.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of elements in one column.",
                  "direction": "in",
                  "name": "kernel_size",
                  "type": "const uint16_t"
                },
                {
                  "description": "per output channel bias. Range : int32",
                  "direction": "in",
                  "name": "output_bias",
                  "type": "const int32_t *const"
                },
                {
                  "description": "pointer to output",
                  "direction": "out",
                  "name": "out",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the two\n\n1. The incremented output pointer for a successful operation or\n2. NULL if implementation is not available."
                }
              ],
              "signature": "int8_t * arm_nn_depthwise_conv_s8_core(\n    const int8_t *row,\n    const int16_t *col,\n    const uint16_t num_ch,\n    const int32_t *out_shift,\n    const int32_t *out_mult,\n    const int32_t out_offset,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const uint16_t kernel_size,\n    const int32_t *const output_bias,\n    int8_t *out\n)",
              "source": {
                "line": 732,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L732"
              },
              "summary": "Depthwise conv on an im2col buffer where the input channel equals output channel."
            },
            {
              "description": "General Matrix-multiplication function with per-channel requantization.\n\nSupported framework: TensorFlow Lite",
              "examples": [],
              "id": "arm_nn_mat_mult_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_s8",
              "params": [
                {
                  "description": "pointer to row operand",
                  "direction": "in",
                  "name": "input_row",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to col operand",
                  "direction": "in",
                  "name": "input_col",
                  "type": "const int8_t *"
                },
                {
                  "description": "number of rows of input_row",
                  "direction": "in",
                  "name": "output_ch",
                  "type": "const uint16_t"
                },
                {
                  "description": "number of column batches. Range: 1 to 4",
                  "direction": "in",
                  "name": "col_batches",
                  "type": "const uint16_t"
                },
                {
                  "description": "pointer to per output channel requantization shift parameter.",
                  "direction": "in",
                  "name": "output_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "pointer to per output channel requantization multiplier parameter.",
                  "direction": "in",
                  "name": "output_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "output tensor offset.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "input tensor(col) offset.",
                  "direction": "in",
                  "name": "col_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "kernel offset(row). Not used.",
                  "direction": "in",
                  "name": "row_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int16_t"
                },
                {
                  "description": "maximum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int16_t"
                },
                {
                  "description": "number of elements in each row",
                  "direction": "in",
                  "name": "row_len",
                  "type": "const uint16_t"
                },
                {
                  "description": "per output channel bias. Range : int32",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *const"
                },
                {
                  "description": "pointer to output",
                  "direction": "inout",
                  "name": "out",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the two\n\n1. The incremented output pointer for a successful operation or\n2. NULL if implementation is not available."
                }
              ],
              "signature": "int8_t * arm_nn_mat_mult_s8(\n    const int8_t *input_row,\n    const int8_t *input_col,\n    const uint16_t output_ch,\n    const uint16_t col_batches,\n    const int32_t *output_shift,\n    const int32_t *output_mult,\n    const int32_t out_offset,\n    const int32_t col_offset,\n    const int32_t row_offset,\n    const int16_t out_activation_min,\n    const int16_t out_activation_max,\n    const uint16_t row_len,\n    const int32_t *const bias,\n    int8_t *out\n)",
              "source": {
                "line": 766,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L766"
              },
              "summary": "General Matrix-multiplication function with per-channel requantization."
            },
            {
              "description": "Matrix-multiplication function for convolution with per-channel requantization for 16 bits convolution.\n\nThis function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.",
              "examples": [],
              "id": "arm_nn_mat_mult_kernel_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_kernel_s16",
              "params": [
                {
                  "description": "pointer to operand A",
                  "direction": "in",
                  "name": "input_a",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to operand B, always consists of 2 vectors.",
                  "direction": "in",
                  "name": "input_b",
                  "type": "const int16_t *"
                },
                {
                  "description": "number of rows of A",
                  "direction": "in",
                  "name": "output_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to per output channel requantization shift parameter.",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "pointer to per output channel requantization multiplier parameter.",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "minimum value to clamp the output to. Range : int16",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp the output to. Range : int16",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of columns of A",
                  "direction": "in",
                  "name": "num_col_a",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const cmsis_nn_bias_data *const"
                },
                {
                  "description": "pointer to output",
                  "direction": "inout",
                  "name": "out_0",
                  "type": "int16_t *"
                },
                {
                  "description": "Address offset between rows in output.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the two\n\n1. The incremented output pointer for a successful operation or\n2. NULL if implementation is not available."
                }
              ],
              "signature": "int16_t * arm_nn_mat_mult_kernel_s16(\n    const int8_t *input_a,\n    const int16_t *input_b,\n    const int32_t output_ch,\n    const int32_t *out_shift,\n    const int32_t *out_mult,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const int32_t num_col_a,\n    const cmsis_nn_bias_data *const bias_data,\n    int16_t *out_0,\n    const int32_t row_address_offset\n)",
              "source": {
                "line": 804,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L804"
              },
              "summary": "Matrix-multiplication function for convolution with per-channel requantization for 16 bits convolution."
            },
            {
              "description": "General Vector by Matrix multiplication with requantization and storage of result.\n\nPseudo-code *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in 'output'",
              "examples": [],
              "id": "arm_nn_mat_mul_core_1x_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mul_core_1x_s8",
              "params": [
                {
                  "description": "number of row elements",
                  "direction": "in",
                  "name": "row_elements",
                  "type": "int32_t"
                },
                {
                  "description": "number of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch",
                  "direction": "in",
                  "name": "skipped_row_elements",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to row operand",
                  "direction": "in",
                  "name": "row_base_ref",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to col operand",
                  "direction": "in",
                  "name": "col_base_ref",
                  "type": "const int8_t *"
                },
                {
                  "description": "Number of output channels",
                  "direction": "in",
                  "name": "out_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to convolution parameters like offsets and activation values",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Pointer to per-channel quantization parameters",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Pointer to optional per-channel bias",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to output where int8 results are stored.",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mul_core_1x_s8(\n    int32_t row_elements,\n    const int32_t skipped_row_elements,\n    const int8_t *row_base_ref,\n    const int8_t *col_base_ref,\n    const int32_t out_ch,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const int32_t *bias,\n    int8_t *output\n)",
              "source": {
                "line": 843,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L843"
              },
              "summary": "General Vector by Matrix multiplication with requantization and storage of result."
            },
            {
              "description": "General Vector by Matrix multiplication with requantization, storage of result and int4 weights packed into an int8 buffer.\n\nPseudo-code as int8 example. Int4 filter data will be unpacked. *output = 0 sum_col = 0 for (j = 0; j < out_ch; j++) for (i = 0; i < row_elements; i++) *output += row_base_ref[i] * col_base_ref[i] sum_col += col_base_ref[i] scale sum_col using quant_params and bias store result in 'output'",
              "examples": [],
              "id": "arm_nn_mat_mul_core_1x_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mul_core_1x_s4",
              "params": [
                {
                  "description": "number of row elements",
                  "direction": "in",
                  "name": "row_elements",
                  "type": "int32_t"
                },
                {
                  "description": "number of row elements skipped due to padding. row_elements + skipped_row_elements = (kernel_x * kernel_y) * input_ch",
                  "direction": "in",
                  "name": "skipped_row_elements",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to row operand",
                  "direction": "in",
                  "name": "row_base_ref",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to col operand as packed int4",
                  "direction": "in",
                  "name": "col_base_ref",
                  "type": "const int8_t *"
                },
                {
                  "description": "Number of output channels",
                  "direction": "in",
                  "name": "out_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to convolution parameters like offsets and activation values",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Pointer to per-channel quantization parameters",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Pointer to optional per-channel bias",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to output where int8 results are stored.",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function performs matrix(row_base_ref) multiplication with vector(col_base_ref) and scaled result is stored in memory."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mul_core_1x_s4(\n    int32_t row_elements,\n    const int32_t skipped_row_elements,\n    const int8_t *row_base_ref,\n    const int8_t *col_base_ref,\n    const int32_t out_ch,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const int32_t *bias,\n    int8_t *output\n)",
              "source": {
                "line": 881,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L881"
              },
              "summary": "General Vector by Matrix multiplication with requantization, storage of result and int4 weights packed into an int8 buffer."
            },
            {
              "description": "Matrix-multiplication with requantization & activation function for four rows and one column.\n\nCompliant to TFLM int8 specification. MVE implementation only",
              "examples": [],
              "id": "arm_nn_mat_mul_core_4x_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mul_core_4x_s8",
              "params": [
                {
                  "description": "number of row elements",
                  "direction": "in",
                  "name": "row_elements",
                  "type": "const int32_t"
                },
                {
                  "description": "offset between rows. Can be the same as row_elements. For e.g, in a 1x1 conv scenario with stride as 1.",
                  "direction": "in",
                  "name": "offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to row operand",
                  "direction": "in",
                  "name": "row_base",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to col operand",
                  "direction": "in",
                  "name": "col_base",
                  "type": "const int8_t *"
                },
                {
                  "description": "Number of output channels",
                  "direction": "in",
                  "name": "out_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "Pointer to convolution parameters like offsets and activation values",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Pointer to per-channel quantization parameters",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Pointer to per-channel bias",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to output where int8 results are stored.",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns the updated output pointer or NULL if implementation is not available."
                }
              ],
              "signature": "int8_t * arm_nn_mat_mul_core_4x_s8(\n    const int32_t row_elements,\n    const int32_t offset,\n    const int8_t *row_base,\n    const int8_t *col_base,\n    const int32_t out_ch,\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const int32_t *bias,\n    int8_t *output\n)",
              "source": {
                "line": 908,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L908"
              },
              "summary": "Matrix-multiplication with requantization & activation function for four rows and one column."
            },
            {
              "description": "General Matrix-multiplication function with per-channel requantization. This function assumes:\n\n- LHS input matrix NOT transposed (nt)\n- RHS input matrix transposed (t)\n- RHS is int8 packed with 2x int4\n- LHS is int8\n\n:::note\nThis operation also performs the broadcast bias addition before the requantization\n\n:::",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_t_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_t_s4",
              "params": [
                {
                  "description": "Pointer to the LHS input matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the RHS input matrix",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the output matrix with \"m\" rows and \"n\" columns",
                  "direction": "out",
                  "name": "dst",
                  "type": "int8_t *"
                },
                {
                  "description": "Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_multipliers",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_shifts",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of LHS input rows",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of RHS input rows",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of LHS/RHS input columns",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be applied to the LHS input value",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be applied the output result",
                  "direction": "in",
                  "name": "dst_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp down the output. Range : int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp up the output. Range : int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Column offset between subsequent lhs_rows",
                  "direction": "in",
                  "name": "lhs_cols_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s4(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t *bias,\n    int8_t *dst,\n    const int32_t *dst_multipliers,\n    const int32_t *dst_shifts,\n    const int32_t lhs_rows,\n    const int32_t rhs_rows,\n    const int32_t rhs_cols,\n    const int32_t lhs_offset,\n    const int32_t dst_offset,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const int32_t lhs_cols_offset\n)",
              "source": {
                "line": 950,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L950"
              },
              "summary": "General Matrix-multiplication function with per-channel requantization."
            },
            {
              "description": "General Matrix-multiplication function with per-channel requantization. This function assumes:\n\n- LHS input matrix NOT transposed (nt)\n- RHS input matrix transposed (t)\n- RHS is int8 packed with 2x int4\n- LHS is int8\n- LHS/RHS input columns must be even numbered\n- LHS must be interleaved. Compare to arm_nn_mat_mult_nt_t_s4 where LHS is not interleaved.\n\n:::note\nThis operation also performs the broadcast bias addition before the requantization\n\n:::",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_interleaved_t_even_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_interleaved_t_even_s4",
              "params": [
                {
                  "description": "Pointer to the LHS input matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the RHS input matrix",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the output matrix with \"m\" rows and \"n\" columns",
                  "direction": "out",
                  "name": "dst",
                  "type": "int8_t *"
                },
                {
                  "description": "Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_multipliers",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_shifts",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of LHS input rows",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of RHS input rows",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of LHS/RHS input columns. Note this must be even.",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be applied to the LHS input value",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be applied the output result",
                  "direction": "in",
                  "name": "dst_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp down the output. Range : int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp up the output. Range : int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Column offset between subsequent lhs_rows",
                  "direction": "in",
                  "name": "lhs_cols_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_interleaved_t_even_s4(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t *bias,\n    int8_t *dst,\n    const int32_t *dst_multipliers,\n    const int32_t *dst_shifts,\n    const int32_t lhs_rows,\n    const int32_t rhs_rows,\n    const int32_t rhs_cols,\n    const int32_t lhs_offset,\n    const int32_t dst_offset,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const int32_t lhs_cols_offset\n)",
              "source": {
                "line": 999,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L999"
              },
              "summary": "General Matrix-multiplication function with per-channel requantization."
            },
            {
              "description": "General Matrix-multiplication function with per-channel requantization. This function assumes:\n\n- LHS input matrix NOT transposed (nt)\n- RHS input matrix transposed (t)\n\n:::note\nThis operation also performs the broadcast bias addition before the requantization\n\n:::",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_t_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_t_s8",
              "params": [
                {
                  "description": "Pointer to the weight sum multiplied by lhs_offset and summed bias buffer",
                  "direction": "in",
                  "name": "weight_sum_buf",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the LHS input matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the RHS input matrix",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the output matrix with \"m\" rows and \"n\" columns",
                  "direction": "out",
                  "name": "dst",
                  "type": "int8_t *"
                },
                {
                  "description": "Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_multipliers",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_shifts",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of LHS input rows",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of RHS input rows",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of LHS/RHS input columns",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be applied to the LHS input value",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be applied the output result",
                  "direction": "in",
                  "name": "dst_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp down the output. Range : int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp up the output. Range : int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Address offset between rows in output. NOTE: Only used for MVEI extension.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Column offset between subsequent lhs_rows",
                  "direction": "in",
                  "name": "lhs_cols_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s8(\n    const int32_t *weight_sum_buf,\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t *bias,\n    int8_t *dst,\n    const int32_t *dst_multipliers,\n    const int32_t *dst_shifts,\n    const int32_t lhs_rows,\n    const int32_t rhs_rows,\n    const int32_t rhs_cols,\n    const int32_t lhs_offset,\n    const int32_t dst_offset,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const int32_t row_address_offset,\n    const int32_t lhs_cols_offset\n)",
              "source": {
                "line": 1046,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1046"
              },
              "summary": "General Matrix-multiplication function with per-channel requantization."
            },
            {
              "description": "General Matrix-multiplication function with per-channel requantization. Output is calculated with multiple channels in parallel, rather than multiple output indices in a single channel This function assumes:\n\n- LHS input matrix NOT transposed (nt)\n- RHS input matrix transposed (t)\n\n:::note\nThis operation also performs the broadcast bias addition before the requantization\n\n:::",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_t_1x1_out_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_t_1x1_out_s8",
              "params": [
                {
                  "description": "Pointer to the weight sum multiplied by lhs_offset and summed bias buffer",
                  "direction": "in",
                  "name": "weight_sum_buf",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the LHS input matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the RHS input matrix",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the bias vector. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the output matrix with \"m\" rows and \"n\" columns",
                  "direction": "out",
                  "name": "dst",
                  "type": "int8_t *"
                },
                {
                  "description": "Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_multipliers",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_shifts",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of LHS input rows",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of RHS input rows",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of LHS/RHS input columns",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be applied to the LHS input value",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be applied the output result",
                  "direction": "in",
                  "name": "dst_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp down the output. Range : int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp up the output. Range : int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Address offset between rows in output. NOTE: Only used for MVEI extension.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Column offset between subsequent lhs_rows",
                  "direction": "in",
                  "name": "lhs_cols_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_t_1x1_out_s8(\n    const int32_t *weight_sum_buf,\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t *bias,\n    int8_t *dst,\n    const int32_t *dst_multipliers,\n    const int32_t *dst_shifts,\n    const int32_t lhs_rows,\n    const int32_t rhs_rows,\n    const int32_t rhs_cols,\n    const int32_t lhs_offset,\n    const int32_t dst_offset,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const int32_t row_address_offset,\n    const int32_t lhs_cols_offset\n)",
              "source": {
                "line": 1097,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1097"
              },
              "summary": "General Matrix-multiplication function with per-channel requantization."
            },
            {
              "description": "General Matrix-multiplication function with per-channel requantization and int16 input (LHS) and output. This function assumes:\n\n- LHS input matrix NOT transposed (nt)\n- RHS input matrix transposed (t)\n\n:::note\nThis operation also performs the broadcast bias addition before the requantization\n\n:::\n\nMVE implementation only.",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_t_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_t_s16",
              "params": [
                {
                  "description": "Pointer to the LHS input matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int16_t *"
                },
                {
                  "description": "Pointer to the RHS input matrix",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to struct with bias vector. The length of this vector is equal to the number of output columns (or RHS input rows). The vector can be int32 or int64 indicated by a flag in the struct.",
                  "direction": "in",
                  "name": "bias_data",
                  "type": "const cmsis_nn_bias_data *"
                },
                {
                  "description": "Pointer to the output matrix with \"m\" rows and \"n\" columns",
                  "direction": "out",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "Pointer to the multipliers vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_multipliers",
                  "type": "const int32_t *"
                },
                {
                  "description": "Pointer to the shifts vector needed for the per-channel requantization. The length of this vector is equal to the number of output columns (or RHS input rows)",
                  "direction": "in",
                  "name": "dst_shifts",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of LHS input rows",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of RHS input rows",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of LHS/RHS input columns",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp down the output. Range : int16",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp up the output. Range : int16",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Address offset between rows in output. NOTE: Only used for MVEI extension.",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` or `ARM_CMSIS_NN_NO_IMPL_ERROR` if not for MVE |---row_address_offset---| |____rhs_rows__________________|\n\n|  |  |\n| --- | --- |\n|  |  |\n|  |  |\n\n| | | lhs_rows\n\n|  |  |\n| --- | --- |\n| _______________ | ______________ |"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s16(\n    const int16_t *lhs,\n    const int8_t *rhs,\n    const cmsis_nn_bias_data *bias_data,\n    int16_t *dst,\n    const int32_t *dst_multipliers,\n    const int32_t *dst_shifts,\n    const int32_t lhs_rows,\n    const int32_t rhs_rows,\n    const int32_t rhs_cols,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const int32_t row_address_offset\n)",
              "source": {
                "line": 1155,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1155"
              },
              "summary": "General Matrix-multiplication function with per-channel requantization and int16 input (LHS) and output."
            },
            {
              "description": "General Matrix-multiplication function with int8 input and int32 output. This function assumes:\n\n- LHS input matrix NOT transposed (nt)\n- RHS input matrix transposed (t)\n\n:::note\nDst/output buffer must be zeroed out before calling this function.\n\n:::",
              "examples": [],
              "id": "arm_nn_mat_mult_nt_t_s8_s32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_nt_t_s8_s32",
              "params": [
                {
                  "description": "Pointer to the LHS input matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the RHS input matrix",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Pointer to the output matrix with \"m\" rows and \"n\" columns. Accumulated into, so it must be zeroed by the caller before the call",
                  "direction": "inout",
                  "name": "dst",
                  "type": "int32_t *"
                },
                {
                  "description": "Number of LHS input rows",
                  "direction": "in",
                  "name": "lhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of LHS input columns/RHS input rows",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of RHS input columns",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be applied to the LHS input value",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset between subsequent output results",
                  "direction": "in",
                  "name": "dst_idx_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_mat_mult_nt_t_s8_s32(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    int32_t *dst,\n    const int32_t lhs_rows,\n    const int32_t rhs_rows,\n    const int32_t rhs_cols,\n    const int32_t lhs_offset,\n    const int32_t dst_idx_offset\n)",
              "source": {
                "line": 1189,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1189"
              },
              "summary": "General Matrix-multiplication function with int8 input and int32 output."
            },
            {
              "description": "s4 Vector by Matrix (transposed) multiplication",
              "examples": [],
              "id": "arm_nn_vec_mat_mult_t_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_vec_mat_mult_t_s4",
              "params": [
                {
                  "description": "Input left-hand side vector",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "packed_rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input bias",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output vector",
                  "direction": "out",
                  "name": "dst",
                  "type": "int8_t *"
                },
                {
                  "description": "Offset to be added to the input values of the left-hand side vector. Range: -127 to 128",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be added to the output values. Range: -127 to 128",
                  "direction": "in",
                  "name": "dst_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "dst_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "dst_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of columns in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of rows in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s4(\n    const int8_t *lhs,\n    const int8_t *packed_rhs,\n    const int32_t *bias,\n    int8_t *dst,\n    const int32_t lhs_offset,\n    const int32_t dst_offset,\n    const int32_t dst_multiplier,\n    const int32_t dst_shift,\n    const int32_t rhs_cols,\n    const int32_t rhs_rows,\n    const int32_t activation_min,\n    const int32_t activation_max\n)",
              "source": {
                "line": 1218,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1218"
              },
              "summary": "s4 Vector by Matrix (transposed) multiplication"
            },
            {
              "description": "s8 Vector by Matrix (transposed) multiplication",
              "examples": [],
              "id": "arm_nn_vec_mat_mult_t_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_vec_mat_mult_t_s8",
              "params": [
                {
                  "description": "Input left-hand side vector",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Kernel sums of the kernels (rhs). See arm_vector_sum_s8 for more info.",
                  "direction": "in",
                  "name": "kernel_sum",
                  "type": "const int32_t *"
                },
                {
                  "description": "Input bias",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output vector",
                  "direction": "out",
                  "name": "dst",
                  "type": "int8_t *"
                },
                {
                  "description": "Offset to be added to the input values of the left-hand side vector. Range: -127 to 128",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be added to the output values. Range: -127 to 128",
                  "direction": "in",
                  "name": "dst_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "dst_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "dst_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of columns in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of rows in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Memory position offset for dst. First output is stored at 'dst', the second at 'dst + address_offset' and so on. Default value is typically 1.",
                  "direction": "in",
                  "name": "address_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be added to the input values of the right-hand side vector. Range: -127 to 128",
                  "direction": "in",
                  "name": "rhs_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s8(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t *kernel_sum,\n    const int32_t *bias,\n    int8_t *dst,\n    const int32_t lhs_offset,\n    const int32_t dst_offset,\n    const int32_t dst_multiplier,\n    const int32_t dst_shift,\n    const int32_t rhs_cols,\n    const int32_t rhs_rows,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const int32_t address_offset,\n    const int32_t rhs_offset\n)",
              "source": {
                "line": 1256,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1256"
              },
              "summary": "s8 Vector by Matrix (transposed) multiplication"
            },
            {
              "description": "s8 Vector by Matrix (transposed) multiplication using per channel quantization for output",
              "examples": [],
              "id": "arm_nn_vec_mat_mult_t_per_ch_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_vec_mat_mult_t_per_ch_s8",
              "params": [
                {
                  "description": "Input left-hand side vector",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Kernel sums of the kernels (rhs). See arm_vector_sum_s8 for more info.",
                  "direction": "in",
                  "name": "kernel_sum",
                  "type": "const int32_t *"
                },
                {
                  "description": "Input bias",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output vector",
                  "direction": "out",
                  "name": "dst",
                  "type": "int8_t *"
                },
                {
                  "description": "Offset to be added to the input values of the left-hand side vector. Range: -127 to 128",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be added to the output values. Range: -127 to 128",
                  "direction": "in",
                  "name": "dst_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multipliers",
                  "direction": "in",
                  "name": "dst_multiplier",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output shifts",
                  "direction": "in",
                  "name": "dst_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of columns in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of rows in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Memory position offset for dst. First output is stored at 'dst', the second at 'dst + address_offset' and so on. Default value is typically 1.",
                  "direction": "in",
                  "name": "address_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset to be added to the input values of the right-hand side vector. Range: -127 to 128",
                  "direction": "in",
                  "name": "rhs_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_vec_mat_mult_t_per_ch_s8(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t *kernel_sum,\n    const int32_t *bias,\n    int8_t *dst,\n    const int32_t lhs_offset,\n    const int32_t dst_offset,\n    const int32_t *dst_multiplier,\n    const int32_t *dst_shift,\n    const int32_t rhs_cols,\n    const int32_t rhs_rows,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const int32_t address_offset,\n    const int32_t rhs_offset\n)",
              "source": {
                "line": 1297,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1297"
              },
              "summary": "s8 Vector by Matrix (transposed) multiplication using per channel quantization for output"
            },
            {
              "description": "s16 Vector by s8 Matrix (transposed) multiplication",
              "examples": [],
              "id": "arm_nn_vec_mat_mult_t_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_vec_mat_mult_t_s16",
              "params": [
                {
                  "description": "Input left-hand side vector",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input bias",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output vector",
                  "direction": "out",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "dst_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "dst_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of columns in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of rows in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int16",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int16",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s16(\n    const int16_t *lhs,\n    const int8_t *rhs,\n    const int64_t *bias,\n    int16_t *dst,\n    const int32_t dst_multiplier,\n    const int32_t dst_shift,\n    const int32_t rhs_cols,\n    const int32_t rhs_rows,\n    const int32_t activation_min,\n    const int32_t activation_max\n)",
              "source": {
                "line": 1330,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1330"
              },
              "summary": "s16 Vector by s8 Matrix (transposed) multiplication"
            },
            {
              "description": "s16 vector(lhs) by s8 matrix (transposed) multiplication and per channel quant output",
              "examples": [],
              "id": "arm_nn_vec_mat_mult_t_per_ch_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_vec_mat_mult_t_per_ch_s16",
              "params": [
                {
                  "description": "Input left-hand side vector",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input bias",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output vector",
                  "direction": "out",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "Per channel output multiplier. Length of vector is equal to rhs_rows",
                  "direction": "in",
                  "name": "dst_multiplier",
                  "type": "const int32_t *"
                },
                {
                  "description": "Per channel output shift. Length of vector is equal to rhs_rows",
                  "direction": "in",
                  "name": "dst_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of columns in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of rows in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int16",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int16",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_vec_mat_mult_t_per_ch_s16(\n    const int16_t *lhs,\n    const int8_t *rhs,\n    const int64_t *bias,\n    int16_t *dst,\n    const int32_t *dst_multiplier,\n    const int32_t *dst_shift,\n    const int32_t rhs_cols,\n    const int32_t rhs_rows,\n    const int32_t activation_min,\n    const int32_t activation_max\n)",
              "source": {
                "line": 1358,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1358"
              },
              "summary": "s16 vector(lhs) by s8 matrix (transposed) multiplication and per channel quant output"
            },
            {
              "description": "s16 Vector by s16 Matrix (transposed) multiplication",
              "examples": [],
              "id": "arm_nn_vec_mat_mult_t_s16_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_vec_mat_mult_t_s16_s16",
              "params": [
                {
                  "description": "Input left-hand side vector",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input bias",
                  "direction": "in",
                  "name": "bias",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output vector",
                  "direction": "out",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "dst_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "dst_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of columns in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of rows in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int16",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int16",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_vec_mat_mult_t_s16_s16(\n    const int16_t *lhs,\n    const int16_t *rhs,\n    const int64_t *bias,\n    int16_t *dst,\n    const int32_t dst_multiplier,\n    const int32_t dst_shift,\n    const int32_t rhs_cols,\n    const int32_t rhs_rows,\n    const int32_t activation_min,\n    const int32_t activation_max\n)",
              "source": {
                "line": 1386,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1386"
              },
              "summary": "s16 Vector by s16 Matrix (transposed) multiplication"
            },
            {
              "description": "s8 Vector by Matrix (transposed) multiplication with s16 output",
              "examples": [],
              "id": "arm_nn_vec_mat_mult_t_svdf_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_vec_mat_mult_t_svdf_s8",
              "params": [
                {
                  "description": "Input left-hand side vector",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Output vector",
                  "direction": "out",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "Offset to be added to the input values of the left-hand side vector. Range: -127 to 128",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Address offset for dst. First output is stored at 'dst', the second at 'dst + scatter_offset' and so on.",
                  "direction": "in",
                  "name": "scatter_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Output multiplier",
                  "direction": "in",
                  "name": "dst_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Output shift",
                  "direction": "in",
                  "name": "dst_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of columns in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of rows in the right-hand side input matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int16",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int16",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_vec_mat_mult_t_svdf_s8(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    int16_t *dst,\n    const int32_t lhs_offset,\n    const int32_t scatter_offset,\n    const int32_t dst_multiplier,\n    const int32_t dst_shift,\n    const int32_t rhs_cols,\n    const int32_t rhs_rows,\n    const int32_t activation_min,\n    const int32_t activation_max\n)",
              "source": {
                "line": 1417,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1417"
              },
              "summary": "s8 Vector by Matrix (transposed) multiplication with s16 output"
            },
            {
              "description": "Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in padded cases where the padding is -lhs_offset(Range: int8). Dimensions are the same for lhs and rhs.\n\n:::note\nTail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `active_ch`.\n\n:::",
              "examples": [],
              "id": "arm_nn_depthwise_conv_nt_t_padded_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv_nt_t_padded_s8",
              "params": [
                {
                  "description": "Input left-hand side matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "LHS matrix offset(input offset). Range: -127 to 128",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Subset of total_ch processed",
                  "direction": "in",
                  "name": "active_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of channels in LHS/RHS",
                  "direction": "in",
                  "name": "total_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "Per channel output shift. Length of vector is equal to number of channels",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "Per channel output multiplier. Length of vector is equal to number of channels",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "Offset to be added to the output values. Range: -127 to 128",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "(row_dimension * col_dimension) of LHS/RHS matrix",
                  "direction": "in",
                  "name": "row_x_col",
                  "type": "const uint16_t"
                },
                {
                  "description": "Per channel output bias. Length of vector is equal to number of channels",
                  "direction": "in",
                  "name": "output_bias",
                  "type": "const int32_t *const"
                },
                {
                  "description": "Output pointer",
                  "direction": "out",
                  "name": "out",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` if an implementation is available or `ARM_CMSIS_NN_NO_IMPL_ERROR` otherwise"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_padded_s8(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t lhs_offset,\n    const int32_t active_ch,\n    const int32_t total_ch,\n    const int32_t *out_shift,\n    const int32_t *out_mult,\n    const int32_t out_offset,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const uint16_t row_x_col,\n    const int32_t *const output_bias,\n    int8_t *out\n)",
              "source": {
                "line": 1453,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1453"
              },
              "summary": "Depthwise convolution of transposed rhs matrix with 4 lhs matrices."
            },
            {
              "description": "Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. Dimensions are the same for lhs and rhs.\n\n:::note\nTail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `active_ch`.\n\n:::",
              "examples": [],
              "id": "arm_nn_depthwise_conv_nt_t_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv_nt_t_s8",
              "params": [
                {
                  "description": "Pointer to the weight sum multiplied by lhs_offset and summed bias buffer",
                  "direction": "in",
                  "name": "weight_sum_buf",
                  "type": "const int32_t *"
                },
                {
                  "description": "Input left-hand side matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "LHS matrix offset(input offset). Range: -127 to 128",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Subset of total_ch processed",
                  "direction": "in",
                  "name": "active_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of channels in LHS/RHS",
                  "direction": "in",
                  "name": "total_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "Per channel output shift. Length of vector is equal to number of channels.",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "Per channel output multiplier. Length of vector is equal to number of channels.",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "Offset to be added to the output values. Range: -127 to 128",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "(row_dimension * col_dimension) of LHS/RHS matrix",
                  "direction": "in",
                  "name": "row_x_col",
                  "type": "const uint16_t"
                },
                {
                  "description": "Per channel output bias. Length of vector is equal to number of channels.",
                  "direction": "in",
                  "name": "output_bias",
                  "type": "const int32_t *const"
                },
                {
                  "description": "Output pointer",
                  "direction": "out",
                  "name": "out",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS` if an implementation is available or `ARM_CMSIS_NN_NO_IMPL_ERROR` otherwise"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_s8(\n    const int32_t *weight_sum_buf,\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t lhs_offset,\n    const int32_t active_ch,\n    const int32_t total_ch,\n    const int32_t *out_shift,\n    const int32_t *out_mult,\n    const int32_t out_offset,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const uint16_t row_x_col,\n    const int32_t *const output_bias,\n    int8_t *out\n)",
              "source": {
                "line": 1492,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1492"
              },
              "summary": "Depthwise convolution of transposed rhs matrix with 4 lhs matrices."
            },
            {
              "description": "Necessary conditions of the planar rule that are cheap to test inline: at most 32 channels and stride 1. A caller can skip `arm_nn_depthwise_conv_s8_planar()` for layers that fail them without changing which layers it takes.",
              "examples": [],
              "id": "arm_nn_depthwise_conv_s8_planar_candidate",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv_s8_planar_candidate",
              "params": [
                {
                  "description": "Depthwise convolution parameters",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [1, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "1 when the layer may take the planar path, 0 when it cannot."
                }
              ],
              "signature": "static int32_t arm_nn_depthwise_conv_s8_planar_candidate(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims\n)",
              "source": {
                "line": 1517,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1517"
              },
              "summary": "Necessary conditions of the planar rule that are cheap to test inline: at most 32 channels and stride 1."
            },
            {
              "description": "The gate of `arm_convolve_s8_small_cin()`: upscale_dims NULL, input depth 1 to 3 with filter depth equal to it, dilation 1, a kernel of at least 1x1 with kernel width x depth at most 16 and at most 48 values, and a positive multiple of 4 output channels. Plain C; it evaluates the same on every build.",
              "examples": [],
              "id": "arm_nn_is_convolve_s8_small_cin",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_is_convolve_s8_small_cin",
              "params": [
                {
                  "description": "Convolution parameters",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, CK]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Upscale tensor dimensions, or NULL",
                  "direction": "in",
                  "name": "upscale_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "1 when the layer is in the gate, 0 otherwise."
                }
              ],
              "signature": "static int32_t arm_nn_is_convolve_s8_small_cin(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    const cmsis_nn_dims *upscale_dims\n)",
              "source": {
                "line": 1536,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1536"
              },
              "summary": "The gate of armconvolves8smallcin(): upscaledims NULL, input depth 1 to 3 with filter depth equal to it, dilation 1, a kernel of at least 1x1 with kernel width…"
            },
            {
              "description": "The gate of `arm_convolve_s8_3x3_c16_s1()`: upscale_dims NULL, input and filter depth 16, a 3x3 kernel, and stride and dilation 1. Plain C; it evaluates the same on every build.",
              "examples": [],
              "id": "arm_nn_is_convolve_s8_3x3_c16_s1",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_is_convolve_s8_3x3_c16_s1",
              "params": [
                {
                  "description": "Convolution parameters",
                  "direction": "in",
                  "name": "conv_params",
                  "type": "const cmsis_nn_conv_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, CK]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Upscale tensor dimensions, or NULL",
                  "direction": "in",
                  "name": "upscale_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "1 when the layer is in the gate, 0 otherwise."
                }
              ],
              "signature": "static int32_t arm_nn_is_convolve_s8_3x3_c16_s1(\n    const cmsis_nn_conv_params *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *upscale_dims\n)",
              "source": {
                "line": 1562,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1562"
              },
              "summary": "The gate of armconvolves83x3c16s1(): upscaledims NULL, input and filter depth 16, a 3x3 kernel, and stride and dilation 1."
            },
            {
              "description": "The group check of `arm_convolve_s8()`, for its direct entries: with groups = C_IN / filter C, C_IN or C_OUT is not a multiple of groups. A filter C of zero or above C_IN gives no group count and is not reported.",
              "examples": [],
              "id": "arm_nn_convolve_s8_groups_invalid",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_convolve_s8_groups_invalid",
              "params": [
                {
                  "description": "Input tensor dimensions. Format: [N, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, CK]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [N, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "1 when `arm_convolve_s8()` reports the group count as an argument error, 0 otherwise."
                }
              ],
              "signature": "static int32_t arm_nn_convolve_s8_groups_invalid(\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1582,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1582"
              },
              "summary": "The group check of armconvolves8(), for its direct entries: with groups = CIN / filter C, CIN or COUT is not a multiple of groups."
            },
            {
              "description": "Plane size in bytes that `arm_nn_depthwise_conv_s8_planar()` needs for a layer, or -1 when the layer is not one it takes. The rule is plain C and evaluates the same on every build.",
              "examples": [],
              "id": "arm_nn_depthwise_conv_s8_planar_bytes",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv_s8_planar_bytes",
              "params": [
                {
                  "description": "Depthwise convolution parameters",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [1, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The plane size in bytes, or -1."
                }
              ],
              "signature": "int32_t arm_nn_depthwise_conv_s8_planar_bytes(\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
              "source": {
                "line": 1601,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1601"
              },
              "summary": "Plane size in bytes that armnndepthwiseconvs8planar() needs for a layer, or -1 when the layer is not one it takes."
            },
            {
              "description": "s8 depthwise convolution with channel multiplier 1 and stride 1, vectorized across the output pixels of one channel plane instead of across channels. It serves the few-channel and 1xk layers of `arm_depthwise_conv_s8_opt()`, with the same scratch buffer and weight sums.",
              "examples": [],
              "id": "arm_nn_depthwise_conv_s8_planar",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv_s8_planar",
              "params": [
                {
                  "description": "Scratch buffer of `arm_depthwise_conv_s8_opt_get_buffer_size()` bytes",
                  "direction": "inout",
                  "name": "ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Per-channel weight sums from `arm_depthwise_convolve_weight_sum()`, bias included",
                  "direction": "in",
                  "name": "weight_sum_ctx",
                  "type": "const cmsis_nn_context *"
                },
                {
                  "description": "Depthwise convolution parameters",
                  "direction": "in",
                  "name": "dw_conv_params",
                  "type": "const cmsis_nn_dw_conv_params *"
                },
                {
                  "description": "Per-channel quantization parameters",
                  "direction": "in",
                  "name": "quant_params",
                  "type": "const cmsis_nn_per_channel_quant_params *"
                },
                {
                  "description": "Input tensor dimensions. Format: [1, H, W, C_IN]",
                  "direction": "in",
                  "name": "input_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Input data pointer",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Filter tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "filter_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Filter data pointer",
                  "direction": "in",
                  "name": "kernel",
                  "type": "const int8_t *"
                },
                {
                  "description": "Output tensor dimensions. Format: [1, H, W, C_OUT]",
                  "direction": "in",
                  "name": "output_dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Output data pointer",
                  "direction": "out",
                  "name": "output",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "`ARM_CMSIS_NN_SUCCESS` when the layer was computed, or `ARM_CMSIS_NN_NO_IMPL_ERROR` when it is not one this path takes or its plane does not fit in ctx->size (then nothing is written), or MVE is not available."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_depthwise_conv_s8_planar(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *weight_sum_ctx,\n    const cmsis_nn_dw_conv_params *dw_conv_params,\n    const cmsis_nn_per_channel_quant_params *quant_params,\n    const cmsis_nn_dims *input_dims,\n    const int8_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const int8_t *kernel,\n    const cmsis_nn_dims *output_dims,\n    int8_t *output\n)",
              "source": {
                "line": 1626,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1626"
              },
              "summary": "s8 depthwise convolution with channel multiplier 1 and stride 1, vectorized across the output pixels of one channel plane instead of across channels."
            },
            {
              "description": "Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. rhs consists of packed int4 data. Dimensions are the same for lhs and rhs.\n\n:::note\nTail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `active_ch`.\n\n:::",
              "examples": [],
              "id": "arm_nn_depthwise_conv_nt_t_s4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv_nt_t_s4",
              "params": [
                {
                  "description": "Input left-hand side matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed). Consists of int4 data packed in an int8 buffer.",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "LHS matrix offset(input offset). Range: -127 to 128",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Subset of total_ch processed",
                  "direction": "in",
                  "name": "active_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of channels in LHS/RHS",
                  "direction": "in",
                  "name": "total_ch",
                  "type": "const int32_t"
                },
                {
                  "description": "Per channel output shift. Length of vector is equal to number of channels.",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "Per channel output multiplier. Length of vector is equal to number of channels.",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "Offset to be added to the output values. Range: -127 to 128",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "(row_dimension * col_dimension) of LHS/RHS matrix",
                  "direction": "in",
                  "name": "row_x_col",
                  "type": "const uint16_t"
                },
                {
                  "description": "Per channel output bias. Length of vector is equal to number of channels.",
                  "direction": "in",
                  "name": "output_bias",
                  "type": "const int32_t *const"
                },
                {
                  "description": "Output pointer",
                  "direction": "out",
                  "name": "out",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the two\n\n- Updated output pointer if an implementation is available\n- NULL if no implementation is available."
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_depthwise_conv_nt_t_s4(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t lhs_offset,\n    const int32_t active_ch,\n    const int32_t total_ch,\n    const int32_t *out_shift,\n    const int32_t *out_mult,\n    const int32_t out_offset,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const uint16_t row_x_col,\n    const int32_t *const output_bias,\n    int8_t *out\n)",
              "source": {
                "line": 1663,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1663"
              },
              "summary": "Depthwise convolution of transposed rhs matrix with 4 lhs matrices."
            },
            {
              "description": "Depthwise convolution of transposed rhs matrix with 4 lhs matrices. To be used in non-padded cases. Dimensions are the same for lhs and rhs.\n\n:::note\nTail channel loads and stores are predicated, so channel-indexed arrays are not accessed beyond `num_ch`.\n\n:::",
              "examples": [],
              "id": "arm_nn_depthwise_conv_nt_t_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_depthwise_conv_nt_t_s16",
              "params": [
                {
                  "description": "Input left-hand side matrix",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int16_t *"
                },
                {
                  "description": "Input right-hand side matrix (transposed)",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Number of channels in LHS/RHS",
                  "direction": "in",
                  "name": "num_ch",
                  "type": "const uint16_t"
                },
                {
                  "description": "Per channel output shift. Length of vector is equal to number of channels.",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "Per channel output multiplier. Length of vector is equal to number of channels.",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "Minimum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Maximum value to clamp the output to. Range: int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "(row_dimension * col_dimension) of LHS/RHS matrix",
                  "direction": "in",
                  "name": "row_x_col",
                  "type": "const uint16_t"
                },
                {
                  "description": "Per channel output bias. Length of vector is equal to number of channels.",
                  "direction": "in",
                  "name": "output_bias",
                  "type": "const int64_t *const"
                },
                {
                  "description": "Output pointer",
                  "direction": "out",
                  "name": "out",
                  "type": "int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the two\n\n- Updated output pointer if an implementation is available\n- NULL if no implementation is available."
                }
              ],
              "signature": "int16_t * arm_nn_depthwise_conv_nt_t_s16(\n    const int16_t *lhs,\n    const int8_t *rhs,\n    const uint16_t num_ch,\n    const int32_t *out_shift,\n    const int32_t *out_mult,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const uint16_t row_x_col,\n    const int64_t *const output_bias,\n    int16_t *out\n)",
              "source": {
                "line": 1699,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1699"
              },
              "summary": "Depthwise convolution of transposed rhs matrix with 4 lhs matrices."
            },
            {
              "description": "Row of s8 scalars multiplicated with a s8 matrix ad accumulated into a s32 rolling scratch buffer. Helpfunction for transposed convolution.\n\n:::note\nRolling buffer refers to how the function wraps around the scratch buffer, e.g. it starts writing at [output_start + output_index], writes to [output_start + output_max] and then continues at [output_start] again.\n\n:::",
              "examples": [],
              "id": "arm_nn_transpose_conv_row_s8_s32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_transpose_conv_row_s8_s32",
              "params": [
                {
                  "description": "Input left-hand side scalars",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Input right-hand side matrix",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Output buffer start",
                  "direction": "out",
                  "name": "output_start",
                  "type": "int32_t *"
                },
                {
                  "description": "Output buffer current index",
                  "direction": "in",
                  "name": "output_index",
                  "type": "const int32_t"
                },
                {
                  "description": "Output buffer size",
                  "direction": "in",
                  "name": "output_max",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of rows in rhs matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of columns in rhs matrix",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of input channels",
                  "direction": "in",
                  "name": "input_channels",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of output channels",
                  "direction": "in",
                  "name": "output_channels",
                  "type": "const int32_t"
                },
                {
                  "description": "Offset added to lhs before multiplication",
                  "direction": "in",
                  "name": "lhs_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Address offset between each row of data output",
                  "direction": "in",
                  "name": "row_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "Length of lhs scalar row.",
                  "direction": "in",
                  "name": "input_x",
                  "type": "const int32_t"
                },
                {
                  "description": "Address offset between each scalar-matrix multiplication result.",
                  "direction": "in",
                  "name": "stride_x",
                  "type": "const int32_t"
                },
                {
                  "description": "Skip rows on top of the filter, used for padding.",
                  "direction": "in",
                  "name": "skip_row_top",
                  "type": "const int32_t"
                },
                {
                  "description": "Skip rows in the bottom of the filter, used for padding.",
                  "direction": "in",
                  "name": "skip_row_bottom",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_transpose_conv_row_s8_s32(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    int32_t *output_start,\n    const int32_t output_index,\n    const int32_t output_max,\n    const int32_t rhs_rows,\n    const int32_t rhs_cols,\n    const int32_t input_channels,\n    const int32_t output_channels,\n    const int32_t lhs_offset,\n    const int32_t row_offset,\n    const int32_t input_x,\n    const int32_t stride_x,\n    const int32_t skip_row_top,\n    const int32_t skip_row_bottom\n)",
              "source": {
                "line": 1735,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1735"
              },
              "summary": "Row of s8 scalars multiplicated with a s8 matrix ad accumulated into a s32 rolling scratch buffer."
            },
            {
              "description": "Read 2 s16 elements and post increment pointer.",
              "examples": [],
              "id": "arm_nn_read_q15x2_ia",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_read_q15x2_ia",
              "params": [
                {
                  "description": "Pointer to pointer that holds address of input. Advanced past the elements read.",
                  "direction": "inout",
                  "name": "in_q15",
                  "type": "const int16_t **"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "q31 value"
                }
              ],
              "signature": "static int32_t arm_nn_read_q15x2_ia(const int16_t **in_q15)",
              "source": {
                "line": 1756,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1756"
              },
              "summary": "Read 2 s16 elements and post increment pointer."
            },
            {
              "description": "Read 4 s8 from s8 pointer and post increment pointer.",
              "examples": [],
              "id": "arm_nn_read_s8x4_ia",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_read_s8x4_ia",
              "params": [
                {
                  "description": "Pointer to pointer that holds address of input. Advanced past the elements read.",
                  "direction": "inout",
                  "name": "in_s8",
                  "type": "const int8_t **"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "q31 value"
                }
              ],
              "signature": "static int32_t arm_nn_read_s8x4_ia(const int8_t **in_s8)",
              "source": {
                "line": 1771,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1771"
              },
              "summary": "Read 4 s8 from s8 pointer and post increment pointer."
            },
            {
              "description": "Read 2 s8 from s8 pointer and post increment pointer.",
              "examples": [],
              "id": "arm_nn_read_s8x2_ia",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_read_s8x2_ia",
              "params": [
                {
                  "description": "Pointer to pointer that holds address of input. Advanced past the elements read.",
                  "direction": "inout",
                  "name": "in_s8",
                  "type": "const int8_t **"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "q31 value"
                }
              ],
              "signature": "static int32_t arm_nn_read_s8x2_ia(const int8_t **in_s8)",
              "source": {
                "line": 1785,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1785"
              },
              "summary": "Read 2 s8 from s8 pointer and post increment pointer."
            },
            {
              "description": "Read 2 int16 values from int16 pointer.",
              "examples": [],
              "id": "arm_nn_read_s16x2",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_read_s16x2",
              "params": [
                {
                  "description": "pointer to address of input.",
                  "direction": "in",
                  "name": "in",
                  "type": "const int16_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "s32 value"
                }
              ],
              "signature": "static int32_t arm_nn_read_s16x2(const int16_t *in)",
              "source": {
                "line": 1799,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1799"
              },
              "summary": "Read 2 int16 values from int16 pointer."
            },
            {
              "description": "Read 4 s8 values.",
              "examples": [],
              "id": "arm_nn_read_s8x4",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_read_s8x4",
              "params": [
                {
                  "description": "pointer to address of input.",
                  "direction": "in",
                  "name": "in_s8",
                  "type": "const int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "s32 value"
                }
              ],
              "signature": "static int32_t arm_nn_read_s8x4(const int8_t *in_s8)",
              "source": {
                "line": 1812,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1812"
              },
              "summary": "Read 4 s8 values."
            },
            {
              "description": "Read 2 s8 values.",
              "examples": [],
              "id": "arm_nn_read_s8x2",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_read_s8x2",
              "params": [
                {
                  "description": "pointer to address of input.",
                  "direction": "in",
                  "name": "in_s8",
                  "type": "const int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "s32 value"
                }
              ],
              "signature": "static int32_t arm_nn_read_s8x2(const int8_t *in_s8)",
              "source": {
                "line": 1824,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1824"
              },
              "summary": "Read 2 s8 values."
            },
            {
              "description": "Write four s8 to s8 pointer and increment pointer afterwards.",
              "examples": [],
              "id": "arm_nn_write_s8x4_ia",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_write_s8x4_ia",
              "params": [
                {
                  "description": "Double pointer to destination. Advanced past the bytes written.",
                  "direction": "inout",
                  "name": "in",
                  "type": "int8_t **"
                },
                {
                  "description": "Four bytes to copy",
                  "direction": "in",
                  "name": "value",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_nn_write_s8x4_ia(int8_t **in, int32_t value)",
              "source": {
                "line": 1837,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1837"
              },
              "summary": "Write four s8 to s8 pointer and increment pointer afterwards."
            },
            {
              "description": "memset optimized for MVE",
              "examples": [],
              "id": "arm_memset_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memset_s8",
              "params": [
                {
                  "description": "Destination pointer",
                  "direction": "inout",
                  "name": "dst",
                  "type": "int8_t *"
                },
                {
                  "description": "Value to set",
                  "direction": "in",
                  "name": "val",
                  "type": "const int8_t"
                },
                {
                  "description": "Number of bytes to copy.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memset_s8(int8_t *dst, const int8_t val, uint32_t block_size)",
              "source": {
                "line": 1850,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1850"
              },
              "summary": "memset optimized for MVE"
            },
            {
              "description": "memset optimized for MVE for 16-bit data.",
              "examples": [],
              "id": "arm_memset_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memset_s16",
              "params": [
                {
                  "description": "Destination pointer.",
                  "direction": "inout",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "16-bit value to set.",
                  "direction": "in",
                  "name": "val",
                  "type": "const int16_t"
                },
                {
                  "description": "Number of int16_t values to set.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memset_s16(int16_t *dst, const int16_t val, uint32_t block_size)",
              "source": {
                "line": 1873,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L1873"
              },
              "summary": "memset optimized for MVE for 16-bit data."
            },
            {
              "description": "Matrix-multiplication function for convolution with per-channel requantization and 4 bit weights.\n\nThis function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.",
              "examples": [],
              "id": "arm_nn_mat_mult_kernel_s4_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_kernel_s4_s16",
              "params": [
                {
                  "description": "pointer to operand A, int8 packed with 2x int4.",
                  "direction": "in",
                  "name": "input_a",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to operand B, always consists of 2 vectors.",
                  "direction": "in",
                  "name": "input_b",
                  "type": "const int16_t *"
                },
                {
                  "description": "number of rows of A",
                  "direction": "in",
                  "name": "output_ch",
                  "type": "const uint16_t"
                },
                {
                  "description": "pointer to per output channel requantization shift parameter.",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "pointer to per output channel requantization multiplier parameter.",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "output tensor offset.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of columns of A",
                  "direction": "in",
                  "name": "num_col_a",
                  "type": "const int32_t"
                },
                {
                  "description": "per output channel bias. Range : int32",
                  "direction": "in",
                  "name": "output_bias",
                  "type": "const int32_t *const"
                },
                {
                  "description": "pointer to output",
                  "direction": "inout",
                  "name": "out_0",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the two\n\n1. The incremented output pointer for a successful operation or\n2. NULL if implementation is not available."
                }
              ],
              "signature": "int8_t * arm_nn_mat_mult_kernel_s4_s16(\n    const int8_t *input_a,\n    const int16_t *input_b,\n    const uint16_t output_ch,\n    const int32_t *out_shift,\n    const int32_t *out_mult,\n    const int32_t out_offset,\n    const int32_t activation_min,\n    const int32_t activation_max,\n    const int32_t num_col_a,\n    const int32_t *const output_bias,\n    int8_t *out_0\n)",
              "source": {
                "line": 2094,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2094"
              },
              "summary": "Matrix-multiplication function for convolution with per-channel requantization and 4 bit weights."
            },
            {
              "description": "Matrix-multiplication function for convolution with per-channel requantization.\n\nThis function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max. Supported framework: TensorFlow Lite micro.",
              "examples": [],
              "id": "arm_nn_mat_mult_kernel_s8_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_kernel_s8_s16",
              "params": [
                {
                  "description": "pointer to operand A",
                  "direction": "in",
                  "name": "input_a",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to operand B, always consists of 2 vectors.",
                  "direction": "in",
                  "name": "input_b",
                  "type": "const int16_t *"
                },
                {
                  "description": "number of rows of A",
                  "direction": "in",
                  "name": "output_ch",
                  "type": "const uint16_t"
                },
                {
                  "description": "pointer to per output channel requantization shift parameter.",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "pointer to per output channel requantization multiplier parameter.",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "output tensor offset.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int16_t"
                },
                {
                  "description": "maximum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int16_t"
                },
                {
                  "description": "number of columns of A",
                  "direction": "in",
                  "name": "num_col_a",
                  "type": "const int32_t"
                },
                {
                  "description": "number of columns of A aligned by 4",
                  "direction": "in",
                  "name": "aligned_num_col_a",
                  "type": "const int32_t"
                },
                {
                  "description": "per output channel bias. Range : int32",
                  "direction": "in",
                  "name": "output_bias",
                  "type": "const int32_t *const"
                },
                {
                  "description": "pointer to output",
                  "direction": "inout",
                  "name": "out_0",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the two\n\n1. The incremented output pointer for a successful operation or\n2. NULL if implementation is not available."
                }
              ],
              "signature": "int8_t * arm_nn_mat_mult_kernel_s8_s16(\n    const int8_t *input_a,\n    const int16_t *input_b,\n    const uint16_t output_ch,\n    const int32_t *out_shift,\n    const int32_t *out_mult,\n    const int32_t out_offset,\n    const int16_t activation_min,\n    const int16_t activation_max,\n    const int32_t num_col_a,\n    const int32_t aligned_num_col_a,\n    const int32_t *const output_bias,\n    int8_t *out_0\n)",
              "source": {
                "line": 2128,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2128"
              },
              "summary": "Matrix-multiplication function for convolution with per-channel requantization."
            },
            {
              "description": "Matrix-multiplication function for convolution with per-channel requantization, supporting an address offset between rows.\n\nThis function does the matrix multiplication of weight matrix for all output channels with 2 columns from im2col and produces two elements/output_channel. The outputs are clamped in the range provided by activation min and max.\n\nThis function is slighly less performant than arm_nn_mat_mult_kernel_s8_s16, but allows support for grouped convolution. Supported framework: TensorFlow Lite micro.",
              "examples": [],
              "id": "arm_nn_mat_mult_kernel_row_offset_s8_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mat_mult_kernel_row_offset_s8_s16",
              "params": [
                {
                  "description": "pointer to operand A",
                  "direction": "in",
                  "name": "input_a",
                  "type": "const int8_t *"
                },
                {
                  "description": "pointer to operand B, always consists of 2 vectors.",
                  "direction": "in",
                  "name": "input_b",
                  "type": "const int16_t *"
                },
                {
                  "description": "number of rows of A",
                  "direction": "in",
                  "name": "output_ch",
                  "type": "const uint16_t"
                },
                {
                  "description": "pointer to per output channel requantization shift parameter.",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t *"
                },
                {
                  "description": "pointer to per output channel requantization multiplier parameter.",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t *"
                },
                {
                  "description": "output tensor offset.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "activation_min",
                  "type": "const int16_t"
                },
                {
                  "description": "maximum value to clamp the output to. Range : int8",
                  "direction": "in",
                  "name": "activation_max",
                  "type": "const int16_t"
                },
                {
                  "description": "number of columns of A",
                  "direction": "in",
                  "name": "num_col_a",
                  "type": "const int32_t"
                },
                {
                  "description": "number of columns of A aligned by 4",
                  "direction": "in",
                  "name": "aligned_num_col_a",
                  "type": "const int32_t"
                },
                {
                  "description": "per output channel bias. Range : int32",
                  "direction": "in",
                  "name": "output_bias",
                  "type": "const int32_t *const"
                },
                {
                  "description": "address offset between rows in the output",
                  "direction": "in",
                  "name": "row_address_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output",
                  "direction": "inout",
                  "name": "out_0",
                  "type": "int8_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns one of the two\n\n1. The incremented output pointer for a successful operation or\n2. NULL if implementation is not available."
                }
              ],
              "signature": "int8_t * arm_nn_mat_mult_kernel_row_offset_s8_s16(\n    const int8_t *input_a,\n    const int16_t *input_b,\n    const uint16_t output_ch,\n    const int32_t *out_shift,\n    const int32_t *out_mult,\n    const int32_t out_offset,\n    const int16_t activation_min,\n    const int16_t activation_max,\n    const int32_t num_col_a,\n    const int32_t aligned_num_col_a,\n    const int32_t *const output_bias,\n    const int32_t row_address_offset,\n    int8_t *out_0\n)",
              "source": {
                "line": 2168,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2168"
              },
              "summary": "Matrix-multiplication function for convolution with per-channel requantization, supporting an address offset between rows."
            },
            {
              "description": "Common softmax function for s8 input and s8 or s16 output.\n\n:::note\nSupported framework: TensorFlow Lite micro (bit-accurate)\n\n:::",
              "examples": [],
              "id": "arm_nn_softmax_common_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_softmax_common_s8",
              "params": [
                {
                  "description": "Pointer to the input tensor",
                  "direction": "in",
                  "name": "input",
                  "type": "const int8_t *"
                },
                {
                  "description": "Number of rows in the input tensor",
                  "direction": "in",
                  "name": "num_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of elements in each input row",
                  "direction": "in",
                  "name": "row_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization multiplier",
                  "direction": "in",
                  "name": "mult",
                  "type": "const int32_t"
                },
                {
                  "description": "Input quantization shift within the range [0, 31]",
                  "direction": "in",
                  "name": "shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Minimum difference with max in row. Used to check if the quantized exponential operation can be performed",
                  "direction": "in",
                  "name": "diff_min",
                  "type": "const int32_t"
                },
                {
                  "description": "Indicating s8 output if 0 else s16 output",
                  "direction": "in",
                  "name": "int16_output",
                  "type": "const bool"
                },
                {
                  "description": "Pointer to the output tensor",
                  "direction": "out",
                  "name": "output",
                  "type": "void *"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "void arm_nn_softmax_common_s8(\n    const int8_t *input,\n    const int32_t num_rows,\n    const int32_t row_size,\n    const int32_t mult,\n    const int32_t shift,\n    const int32_t diff_min,\n    const bool int16_output,\n    void *output\n)",
              "source": {
                "line": 2197,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2197"
              },
              "summary": "Common softmax function for s8 input and s8 or s16 output."
            },
            {
              "description": "macro for adding rounding offset",
              "examples": [],
              "id": "NN_ROUND",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "NN_ROUND",
              "params": [
                {
                  "description": "",
                  "name": "out_shift"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define NN_ROUND(out_shift) ((0x1 << out_shift) >> 1)",
              "source": {
                "line": 2210,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2210"
              },
              "summary": "macro for adding rounding offset"
            },
            {
              "description": "",
              "examples": [],
              "id": "MUL_SAT",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "MUL_SAT",
              "params": [
                {
                  "description": "",
                  "name": "a"
                },
                {
                  "description": "",
                  "name": "b"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define MUL_SAT(a, b) arm_nn_doubling_high_mult((a), (b))",
              "source": {
                "line": 2216,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2216"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "MUL_SAT_MVE",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "MUL_SAT_MVE",
              "params": [
                {
                  "description": "",
                  "name": "a"
                },
                {
                  "description": "",
                  "name": "b"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define MUL_SAT_MVE(a, b) arm_doubling_high_mult_mve_32x4((a), (b))",
              "source": {
                "line": 2217,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2217"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "MUL_POW2",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "MUL_POW2",
              "params": [
                {
                  "description": "",
                  "name": "a"
                },
                {
                  "description": "",
                  "name": "b"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define MUL_POW2(a, b) arm_nn_mult_by_power_of_two((a), (b))",
              "source": {
                "line": 2218,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2218"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "DIV_POW2",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "DIV_POW2",
              "params": [
                {
                  "description": "",
                  "name": "a"
                },
                {
                  "description": "",
                  "name": "b"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define DIV_POW2(a, b) arm_nn_divide_by_power_of_two((a), (b))",
              "source": {
                "line": 2220,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2220"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "DIV_POW2_MVE",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "DIV_POW2_MVE",
              "params": [
                {
                  "description": "",
                  "name": "a"
                },
                {
                  "description": "",
                  "name": "b"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define DIV_POW2_MVE(a, b) arm_divide_by_power_of_two_mve((a), (b))",
              "source": {
                "line": 2221,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2221"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "EXP_ON_NEG",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "EXP_ON_NEG",
              "params": [
                {
                  "description": "",
                  "name": "x"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define EXP_ON_NEG(x) arm_nn_exp_on_negative_values((x))",
              "source": {
                "line": 2223,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2223"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "ONE_OVER1",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ONE_OVER1",
              "params": [
                {
                  "description": "",
                  "name": "x"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define ONE_OVER1(x) arm_nn_one_over_one_plus_x_for_x_in_0_1((x))",
              "source": {
                "line": 2224,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2224"
              },
              "summary": ""
            },
            {
              "description": "Saturating doubling high multiply. Result matches NEON instruction VQRDMULH.",
              "examples": [],
              "id": "arm_nn_doubling_high_mult",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_doubling_high_mult",
              "params": [
                {
                  "description": "Multiplicand. Range: {NN_Q31_MIN, NN_Q31_MAX}",
                  "direction": "in",
                  "name": "m1",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier. Range: {NN_Q31_MIN, NN_Q31_MAX}",
                  "direction": "in",
                  "name": "m2",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Result of multiplication."
                }
              ],
              "signature": "static int32_t arm_nn_doubling_high_mult(const int32_t m1, const int32_t m2)",
              "source": {
                "line": 2234,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2234"
              },
              "summary": "Saturating doubling high multiply."
            },
            {
              "description": "Doubling high multiply without saturation. This is intended for requantization where the scale is a positive integer.\n\n:::note\nThe result of this matches that of neon instruction VQRDMULH for m1 in range {NN_Q31_MIN, NN_Q31_MAX} and m2 in range {NN_Q31_MIN + 1, NN_Q31_MAX}. Saturation occurs when m1 equals m2 equals NN_Q31_MIN and that is not handled by this function.\n\n:::",
              "examples": [],
              "id": "arm_nn_doubling_high_mult_no_sat",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_doubling_high_mult_no_sat",
              "params": [
                {
                  "description": "Multiplicand. Range: {NN_Q31_MIN, NN_Q31_MAX}",
                  "direction": "in",
                  "name": "m1",
                  "type": "int32_t"
                },
                {
                  "description": "Multiplier Range: {NN_Q31_MIN, NN_Q31_MAX}",
                  "direction": "in",
                  "name": "m2",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Result of multiplication."
                }
              ],
              "signature": "static int32_t arm_nn_doubling_high_mult_no_sat(int32_t m1, int32_t m2)",
              "source": {
                "line": 2272,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2272"
              },
              "summary": "Doubling high multiply without saturation."
            },
            {
              "description": "Rounding divide by power of two.",
              "examples": [],
              "id": "arm_nn_divide_by_power_of_two",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_divide_by_power_of_two",
              "params": [
                {
                  "description": "- Dividend",
                  "direction": "in",
                  "name": "dividend",
                  "type": "const int32_t"
                },
                {
                  "description": "- Divisor = power(2, exponent) Range: [0, 31]",
                  "direction": "in",
                  "name": "exponent",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Rounded result of division. Midpoint is rounded away from zero."
                }
              ],
              "signature": "static int32_t arm_nn_divide_by_power_of_two(const int32_t dividend, const int32_t exponent)",
              "source": {
                "line": 2323,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2323"
              },
              "summary": "Rounding divide by power of two."
            },
            {
              "description": "Rounding divide by power of two for non-negative values.",
              "examples": [],
              "id": "arm_nn_nonneg_divide_by_pot_s32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_nonneg_divide_by_pot_s32",
              "params": [
                {
                  "description": "- Dividend (assumed to be non-negative)",
                  "direction": "in",
                  "name": "dividend",
                  "type": "int32_t"
                },
                {
                  "description": "- Divisor = power(2, exponent) Range: [0, 31]",
                  "direction": "in",
                  "name": "exponent",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Rounded result of division. Midpoint is rounded away from zero."
                }
              ],
              "signature": "static int32_t arm_nn_nonneg_divide_by_pot_s32(int32_t dividend, int32_t exponent)",
              "source": {
                "line": 2378,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2378"
              },
              "summary": "Rounding divide by power of two for non-negative values."
            },
            {
              "description": "Requantize a given value.\n\nEssentially returns (val * multiplier)/(2 ^ shift) with different rounding depending if CMSIS_NN_USE_SINGLE_ROUNDING is defined or not.",
              "examples": [],
              "id": "arm_nn_requantize",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_requantize",
              "params": [
                {
                  "description": "Value to be requantized",
                  "direction": "in",
                  "name": "val",
                  "type": "const int32_t"
                },
                {
                  "description": "Multiplier. Range {NN_Q31_MIN + 1, Q32_MAX}",
                  "direction": "in",
                  "name": "multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift. Range: {-31, 30} Default branch: If shift is positive left shift 'val * multiplier' with shift If shift is negative right shift 'val * multiplier' with abs(shift) Single round branch: Input for total_shift in divide by '2 ^ total_shift'",
                  "direction": "in",
                  "name": "shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Default branch: Returns (val * multiplier) with rounding divided by (2 ^ shift) with rounding Single round branch: Returns (val * multiplier)/(2 ^ (31 - shift)) with rounding"
                }
              ],
              "signature": "static int32_t arm_nn_requantize(const int32_t val, const int32_t multiplier, const int32_t shift)",
              "source": {
                "line": 2416,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2416"
              },
              "summary": "Requantize a given value."
            },
            {
              "description": "Requantize a given 64 bit value.",
              "examples": [],
              "id": "arm_nn_requantize_s64",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_requantize_s64",
              "params": [
                {
                  "description": "Value to be requantized in the range {-(1<<47)} to {(1<<47) - 1}",
                  "direction": "in",
                  "name": "val",
                  "type": "const int64_t"
                },
                {
                  "description": "Reduced multiplier in the range {NN_Q31_MIN + 1, Q32_MAX} to {Q16_MIN + 1, Q16_MAX}",
                  "direction": "in",
                  "name": "reduced_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Left or right shift for 'val * multiplier' in the range {-31} to {7}",
                  "direction": "in",
                  "name": "shift",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Returns (val * multiplier)/(2 ^ shift)"
                }
              ],
              "signature": "static int32_t arm_nn_requantize_s64(const int64_t val, const int32_t reduced_multiplier, const int32_t shift)",
              "source": {
                "line": 2453,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2453"
              },
              "summary": "Requantize a given 64 bit value."
            },
            {
              "description": "Saturating left shift for int16_t.",
              "examples": [],
              "id": "arm_nn_sat_lshift_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_sat_lshift_s16",
              "params": [
                {
                  "description": "value to be shifted",
                  "direction": "in",
                  "name": "x",
                  "type": "int16_t"
                },
                {
                  "description": "Nonpositive values return x; positive values multiply by 2^shift with s16 saturation.",
                  "direction": "in",
                  "name": "shift",
                  "type": "int"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "shifted value"
                }
              ],
              "signature": "static int16_t arm_nn_sat_lshift_s16(int16_t x, int shift)",
              "source": {
                "line": 2471,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2471"
              },
              "summary": "Saturating left shift for int16t."
            },
            {
              "description": "Saturating *Rounding* Doubling High Mul (s16).\n\nMatches NEON SQRDMULH s16",
              "examples": [],
              "id": "arm_nn_sqrdmulh_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_sqrdmulh_s16",
              "params": [
                {
                  "description": "Multiplicand",
                  "direction": "in",
                  "name": "a",
                  "type": "int16_t"
                },
                {
                  "description": "Multiplier",
                  "direction": "in",
                  "name": "b",
                  "type": "int16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Result of multiplication."
                }
              ],
              "signature": "static int16_t arm_nn_sqrdmulh_s16(int16_t a, int16_t b)",
              "source": {
                "line": 2490,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2490"
              },
              "summary": "Saturating Rounding Doubling High Mul (s16)."
            },
            {
              "description": "Saturating **Non-rounded** Doubling High Mul (s16).\n\nMatches NEON SQDMULH s16",
              "examples": [],
              "id": "arm_nn_sqdmulh_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_sqdmulh_s16",
              "params": [
                {
                  "description": "Multiplicand",
                  "direction": "in",
                  "name": "a",
                  "type": "int16_t"
                },
                {
                  "description": "Multiplier",
                  "direction": "in",
                  "name": "b",
                  "type": "int16_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Result of multiplication."
                }
              ],
              "signature": "static int16_t arm_nn_sqdmulh_s16(int16_t a, int16_t b)",
              "source": {
                "line": 2510,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2510"
              },
              "summary": "Saturating Non-rounded Doubling High Mul (s16)."
            },
            {
              "description": "Rounding divide by power of two (s16), midpoint away from zero.\n\nMirrors arm_nn_divide_by_power_of_two() semantics for s16.",
              "examples": [],
              "id": "arm_nn_divide_by_power_of_two_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_divide_by_power_of_two_s16",
              "params": [
                {
                  "description": "Dividend",
                  "direction": "in",
                  "name": "x",
                  "type": "int16_t"
                },
                {
                  "description": "Divisor = power(2, exponent) Range: [0, 15]",
                  "direction": "in",
                  "name": "exponent",
                  "type": "int"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Rounded result of division. Midpoint is rounded away from zero."
                }
              ],
              "signature": "static int16_t arm_nn_divide_by_power_of_two_s16(int16_t x, int exponent)",
              "source": {
                "line": 2530,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2530"
              },
              "summary": "Rounding divide by power of two (s16), midpoint away from zero."
            },
            {
              "description": "memcpy optimized for MVE",
              "examples": [],
              "id": "arm_memcpy_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memcpy_s8",
              "params": [
                {
                  "description": "Destination pointer",
                  "direction": "inout",
                  "name": "dst",
                  "type": "int8_t *"
                },
                {
                  "description": "Source pointer.",
                  "direction": "in",
                  "name": "src",
                  "type": "const int8_t *"
                },
                {
                  "description": "Number of bytes to copy.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memcpy_s8(int8_t *dst, const int8_t *src, uint32_t block_size)",
              "source": {
                "line": 2545,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2545"
              },
              "summary": "memcpy optimized for MVE"
            },
            {
              "description": "memcpy optimized for MVE",
              "examples": [],
              "id": "arm_memcpy_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memcpy_s16",
              "params": [
                {
                  "description": "Destination pointer",
                  "direction": "inout",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "Source pointer.",
                  "direction": "in",
                  "name": "src",
                  "type": "const int16_t *"
                },
                {
                  "description": "Number of values to copy.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memcpy_s16(int16_t *dst, const int16_t *src, uint32_t block_size)",
              "source": {
                "line": 2569,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2569"
              },
              "summary": "memcpy optimized for MVE"
            },
            {
              "description": "memcpy optimized for MVE",
              "examples": [],
              "id": "arm_memcpy_s32",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memcpy_s32",
              "params": [
                {
                  "description": "Destination pointer",
                  "direction": "inout",
                  "name": "dst",
                  "type": "int32_t *"
                },
                {
                  "description": "Source pointer.",
                  "direction": "in",
                  "name": "src",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of values to copy.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memcpy_s32(int32_t *dst, const int32_t *src, uint32_t block_size)",
              "source": {
                "line": 2581,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2581"
              },
              "summary": "memcpy optimized for MVE"
            },
            {
              "description": "memcpy wrapper for int16",
              "examples": [],
              "id": "arm_memcpy_q15",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_memcpy_q15",
              "params": [
                {
                  "description": "Destination pointer",
                  "direction": "inout",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "Source pointer.",
                  "direction": "in",
                  "name": "src",
                  "type": "const int16_t *"
                },
                {
                  "description": "Number of bytes to copy.",
                  "direction": "in",
                  "name": "block_size",
                  "type": "uint32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_memcpy_q15(int16_t *dst, const int16_t *src, uint32_t block_size)",
              "source": {
                "line": 2593,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2593"
              },
              "summary": "memcpy wrapper for int16"
            },
            {
              "description": "Fixed-point exp() of a non-positive value.",
              "examples": [],
              "id": "arm_nn_exp_on_negative_values",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_exp_on_negative_values",
              "params": [
                {
                  "description": "Input in Q5.26 fixed point. Must be less than or equal to 0",
                  "direction": "in",
                  "name": "val",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "exp(val) in Q0.31 fixed point. Returns NN_Q31_MAX when `val` is 0."
                }
              ],
              "signature": "static int32_t arm_nn_exp_on_negative_values(int32_t val)",
              "source": {
                "line": 2865,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2865"
              },
              "summary": "Fixed-point exp() of a non-positive value."
            },
            {
              "description": "",
              "examples": [],
              "id": "SELECT_IF_NON_ZERO",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "SELECT_IF_NON_ZERO",
              "params": [
                {
                  "description": "",
                  "name": "x"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "#define SELECT_IF_NON_ZERO(x) { \\ mask = MASK_IF_NON_ZERO(remainder & (1 << shift++)); \\ result = SELECT_USING_MASK(mask, MUL_SAT(result, x), result); \\ }",
              "source": {
                "line": 2878,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2878"
              },
              "summary": ""
            },
            {
              "description": "Saturating multiply by a power of two.",
              "examples": [],
              "id": "arm_nn_mult_by_power_of_two",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_mult_by_power_of_two",
              "params": [
                {
                  "description": "Value to be multiplied",
                  "direction": "in",
                  "name": "val",
                  "type": "const int32_t"
                },
                {
                  "description": "Exponent. Multiplier = power(2, exp)",
                  "direction": "in",
                  "name": "exp",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "val * 2^exp saturated to the int32 range"
                }
              ],
              "signature": "static int32_t arm_nn_mult_by_power_of_two(const int32_t val, const int32_t exp)",
              "source": {
                "line": 2905,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2905"
              },
              "summary": "Saturating multiply by a power of two."
            },
            {
              "description": "Fixed-point 1 / (1 + x) for x in [0, 1), computed with Newton-Raphson iterations.",
              "examples": [],
              "id": "arm_nn_one_over_one_plus_x_for_x_in_0_1",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_one_over_one_plus_x_for_x_in_0_1",
              "params": [
                {
                  "description": "x in Q0.31 fixed point. Range: [0, NN_Q31_MAX]",
                  "direction": "in",
                  "name": "val",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "1 / (1 + x) in Q0.31 fixed point"
                }
              ],
              "signature": "static int32_t arm_nn_one_over_one_plus_x_for_x_in_0_1(int32_t val)",
              "source": {
                "line": 2920,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2920"
              },
              "summary": "Fixed-point 1 / (1 + x) for x in [0, 1), computed with Newton-Raphson iterations."
            },
            {
              "description": "Write 2 s16 elements and post increment pointer.",
              "examples": [],
              "id": "arm_nn_write_q15x2_ia",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_write_q15x2_ia",
              "params": [
                {
                  "description": "Pointer to pointer that holds address of destination. Advanced past the elements written.",
                  "direction": "inout",
                  "name": "dest_q15",
                  "type": "int16_t **"
                },
                {
                  "description": "Input value to be written.",
                  "direction": "in",
                  "name": "src_q31",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_nn_write_q15x2_ia(int16_t **dest_q15, int32_t src_q31)",
              "source": {
                "line": 2942,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2942"
              },
              "summary": "Write 2 s16 elements and post increment pointer."
            },
            {
              "description": "Write 2 s8 elements and post increment pointer.",
              "examples": [],
              "id": "arm_nn_write_s8x2_ia",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_write_s8x2_ia",
              "params": [
                {
                  "description": "Pointer to pointer that holds address of destination. Advanced past the elements written.",
                  "direction": "inout",
                  "name": "dst",
                  "type": "int8_t **"
                },
                {
                  "description": "Input value to be written.",
                  "direction": "in",
                  "name": "src",
                  "type": "int16_t"
                }
              ],
              "raises": [],
              "returns": [],
              "signature": "static void arm_nn_write_s8x2_ia(int8_t **dst, int16_t src)",
              "source": {
                "line": 2955,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2955"
              },
              "summary": "Write 2 s8 elements and post increment pointer."
            },
            {
              "description": "Get dimension value at specific index.",
              "examples": [],
              "id": "arm_cmsis_nn_dim_at",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_cmsis_nn_dim_at",
              "params": [
                {
                  "description": "Pointer to `cmsis_nn_dims` structure",
                  "direction": "in",
                  "name": "dims",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "Index of dimension to get",
                  "direction": "in",
                  "name": "index",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Dimension value at specified index"
                }
              ],
              "signature": "static int32_t arm_cmsis_nn_dim_at(const cmsis_nn_dims *dims, int32_t index)",
              "source": {
                "line": 2969,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2969"
              },
              "summary": "Get dimension value at specific index."
            },
            {
              "description": "Calculate the product of all dimensions in a shape array.",
              "examples": [],
              "id": "arm_cmsis_nn_shape_product",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_cmsis_nn_shape_product",
              "params": [
                {
                  "description": "Pointer to array containing shape dimensions",
                  "direction": "in",
                  "name": "shape",
                  "type": "const int32_t *"
                },
                {
                  "description": "Number of dimensions in the shape array",
                  "direction": "in",
                  "name": "length",
                  "type": "int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "Product of all dimensions"
                }
              ],
              "signature": "static size_t arm_cmsis_nn_shape_product(const int32_t *shape, int32_t length)",
              "source": {
                "line": 2994,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L2994"
              },
              "summary": "Calculate the product of all dimensions in a shape array."
            },
            {
              "description": "Update LSTM function for an iteration step using s8 input and output, and s16 internally.",
              "examples": [],
              "id": "arm_nn_lstm_step_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_lstm_step_s8",
              "params": [
                {
                  "description": "Data input pointer",
                  "direction": "in",
                  "name": "data_in",
                  "type": "const int8_t *"
                },
                {
                  "description": "Hidden state/ recurrent input pointer",
                  "direction": "in",
                  "name": "hidden_in",
                  "type": "const int8_t *"
                },
                {
                  "description": "Hidden state/ recurrent output pointer",
                  "direction": "out",
                  "name": "hidden_out",
                  "type": "int8_t *"
                },
                {
                  "description": "Struct containg all information about the lstm operator, see arm_nn_types.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params *"
                },
                {
                  "description": "Struct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_lstm_context *"
                },
                {
                  "description": "Number of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_lstm_step_s8(\n    const int8_t *data_in,\n    const int8_t *hidden_in,\n    int8_t *hidden_out,\n    const cmsis_nn_lstm_params *params,\n    cmsis_nn_lstm_context *buffers,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 3022,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3022"
              },
              "summary": "Update LSTM function for an iteration step using s8 input and output, and s16 internally."
            },
            {
              "description": "Update LSTM function for an iteration step using s16 input and output, and s16 internally.",
              "examples": [],
              "id": "arm_nn_lstm_step_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_lstm_step_s16",
              "params": [
                {
                  "description": "Data input pointer",
                  "direction": "in",
                  "name": "data_in",
                  "type": "const int16_t *"
                },
                {
                  "description": "Hidden state/ recurrent input pointer",
                  "direction": "in",
                  "name": "hidden_in",
                  "type": "const int16_t *"
                },
                {
                  "description": "Hidden state/ recurrent output pointer",
                  "direction": "out",
                  "name": "hidden_out",
                  "type": "int16_t *"
                },
                {
                  "description": "Struct containg all information about the lstm operator, see arm_nn_types.",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params *"
                },
                {
                  "description": "Struct containg pointers to all temporary scratch buffers needed for the lstm operator, see arm_nn_types.",
                  "direction": "inout",
                  "name": "buffers",
                  "type": "cmsis_nn_lstm_context *"
                },
                {
                  "description": "Number of timesteps between consecutive batches. E.g for params->timing_major = true, all batches for t=0 are stored sequentially, so batch offset = 1. For params->time major = false, all time steps are stored continously before the next batch, so batch offset = params->time_steps.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_lstm_step_s16(\n    const int16_t *data_in,\n    const int16_t *hidden_in,\n    int16_t *hidden_out,\n    const cmsis_nn_lstm_params *params,\n    cmsis_nn_lstm_context *buffers,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 3046,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3046"
              },
              "summary": "Update LSTM function for an iteration step using s16 input and output, and s16 internally."
            },
            {
              "description": "Updates a LSTM gate for an iteration step of LSTM function, int8x8_16 version.",
              "examples": [],
              "id": "arm_nn_lstm_calculate_gate_s8_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_lstm_calculate_gate_s8_s16",
              "params": [
                {
                  "description": "Data input pointer",
                  "direction": "in",
                  "name": "data_in",
                  "type": "const int8_t *"
                },
                {
                  "description": "Hidden state/ recurrent input pointer",
                  "direction": "in",
                  "name": "hidden_in",
                  "type": "const int8_t *"
                },
                {
                  "description": "Struct containing all information about the gate caluclation, see arm_nn_types.",
                  "direction": "in",
                  "name": "gate_data",
                  "type": "const cmsis_nn_lstm_gate *"
                },
                {
                  "description": "Struct containing all information about the lstm_operation, see arm_nn_types",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params *"
                },
                {
                  "description": "Hidden state/ recurrent output pointer",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of timesteps between consecutive batches, see arm_nn_lstm_step_s8.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_lstm_calculate_gate_s8_s16(\n    const int8_t *data_in,\n    const int8_t *hidden_in,\n    const cmsis_nn_lstm_gate *gate_data,\n    const cmsis_nn_lstm_params *params,\n    int16_t *output,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 3067,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3067"
              },
              "summary": "Updates a LSTM gate for an iteration step of LSTM function, int8x816 version."
            },
            {
              "description": "Updates a LSTM gate for an iteration step of LSTM function, int16x8_16 version.",
              "examples": [],
              "id": "arm_nn_lstm_calculate_gate_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_lstm_calculate_gate_s16",
              "params": [
                {
                  "description": "Data input pointer",
                  "direction": "in",
                  "name": "data_in",
                  "type": "const int16_t *"
                },
                {
                  "description": "Hidden state/ recurrent input pointer",
                  "direction": "in",
                  "name": "hidden_in",
                  "type": "const int16_t *"
                },
                {
                  "description": "Struct containing all information about the gate caluclation, see arm_nn_types.",
                  "direction": "in",
                  "name": "gate_data",
                  "type": "const cmsis_nn_lstm_gate *"
                },
                {
                  "description": "Struct containing all information about the lstm_operation, see arm_nn_types",
                  "direction": "in",
                  "name": "params",
                  "type": "const cmsis_nn_lstm_params *"
                },
                {
                  "description": "Hidden state/ recurrent output pointer",
                  "direction": "out",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "Number of timesteps between consecutive batches, see arm_nn_lstm_step_s16.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_lstm_calculate_gate_s16(\n    const int16_t *data_in,\n    const int16_t *hidden_in,\n    const cmsis_nn_lstm_gate *gate_data,\n    const cmsis_nn_lstm_params *params,\n    int16_t *output,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 3088,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3088"
              },
              "summary": "Updates a LSTM gate for an iteration step of LSTM function, int16x816 version."
            },
            {
              "description": "The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a \"batched\" vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).",
              "examples": [],
              "id": "arm_nn_vec_mat_mul_result_acc_s8_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_vec_mat_mul_result_acc_s8_s16",
              "params": [
                {
                  "description": "Batched vector",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Weights - input matrix (H(Rows)xW(Columns))",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias + lhs_offset * kernel_sum term precalculated into a constant vector.",
                  "direction": "in",
                  "name": "effective_bias",
                  "type": "const int32_t *"
                },
                {
                  "description": "Output",
                  "direction": "out",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "Multiplier for quantization",
                  "direction": "in",
                  "name": "dst_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for quantization",
                  "direction": "in",
                  "name": "dst_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Vector/matarix column length",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Row count of matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Batch size",
                  "direction": "in",
                  "name": "batches",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of timesteps between consecutive batches in input, see arm_nn_lstm_step_s8. Note that the output is always stored with sequential batches.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_vec_mat_mul_result_acc_s8_s16(\n    const int8_t *lhs,\n    const int8_t *rhs,\n    const int32_t *effective_bias,\n    int16_t *dst,\n    const int32_t dst_multiplier,\n    const int32_t dst_shift,\n    const int32_t rhs_cols,\n    const int32_t rhs_rows,\n    const int32_t batches,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 3114,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3114"
              },
              "summary": "The result of the multiplication is accumulated to the passed result buffer."
            },
            {
              "description": "The result of the multiplication is accumulated to the passed result buffer. Multiplies a matrix by a \"batched\" vector (i.e. a matrix with a batch dimension composed by input vectors independent from each other).",
              "examples": [],
              "id": "arm_nn_vec_mat_mul_result_acc_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_vec_mat_mul_result_acc_s16",
              "params": [
                {
                  "description": "Batched vector",
                  "direction": "in",
                  "name": "lhs",
                  "type": "const int16_t *"
                },
                {
                  "description": "Weights - input matrix (H(Rows)xW(Columns))",
                  "direction": "in",
                  "name": "rhs",
                  "type": "const int8_t *"
                },
                {
                  "description": "Bias + lhs_offset * kernel_sum term precalculated into a constant vector.",
                  "direction": "in",
                  "name": "effective_bias",
                  "type": "const int64_t *"
                },
                {
                  "description": "Output",
                  "direction": "out",
                  "name": "dst",
                  "type": "int16_t *"
                },
                {
                  "description": "Multiplier for quantization",
                  "direction": "in",
                  "name": "dst_multiplier",
                  "type": "const int32_t"
                },
                {
                  "description": "Shift for quantization",
                  "direction": "in",
                  "name": "dst_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "Vector/matarix column length",
                  "direction": "in",
                  "name": "rhs_cols",
                  "type": "const int32_t"
                },
                {
                  "description": "Row count of matrix",
                  "direction": "in",
                  "name": "rhs_rows",
                  "type": "const int32_t"
                },
                {
                  "description": "Batch size",
                  "direction": "in",
                  "name": "batches",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of timesteps between consecutive batches in input, see arm_nn_lstm_step_s16. Note that the output is always stored with sequential batches.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns `ARM_CMSIS_NN_SUCCESS`"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_nn_vec_mat_mul_result_acc_s16(\n    const int16_t *lhs,\n    const int8_t *rhs,\n    const int64_t *effective_bias,\n    int16_t *dst,\n    const int32_t dst_multiplier,\n    const int32_t dst_shift,\n    const int32_t rhs_cols,\n    const int32_t rhs_rows,\n    const int32_t batches,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 3144,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3144"
              },
              "summary": "The result of the multiplication is accumulated to the passed result buffer."
            },
            {
              "description": "s16 elementwise multiplication with s8 output\n\nSupported framework: TensorFlow Lite micro",
              "examples": [],
              "id": "arm_elementwise_mul_s16_s8",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_mul_s16_s8",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "inout",
                  "name": "output",
                  "type": "int8_t *"
                },
                {
                  "description": "output offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples per batch",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples per batch",
                  "direction": "in",
                  "name": "batch_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of timesteps between consecutive batches in output, see arm_nn_lstm_step_s8. Note that it is assumed that the input is stored with sequential batches.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_mul_s16_s8(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    int8_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t block_size,\n    const int32_t batch_size,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 3171,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3171"
              },
              "summary": "s16 elementwise multiplication with s8 output"
            },
            {
              "description": "s16 elementwise multiplication with s16 output\n\nSupported framework: TensorFlow Lite micro",
              "examples": [],
              "id": "arm_elementwise_mul_s16_batch_offset",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_mul_s16_batch_offset",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "inout",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples per batch",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples per batch",
                  "direction": "in",
                  "name": "batch_size",
                  "type": "const int32_t"
                },
                {
                  "description": "Number of timesteps between consecutive batches in output, see arm_nn_lstm_step_s16. Note that it is assumed that the input is stored with sequential batches.",
                  "direction": "in",
                  "name": "batch_offset",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_mul_s16_batch_offset(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t block_size,\n    const int32_t batch_size,\n    const int32_t batch_offset\n)",
              "source": {
                "line": 3197,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3197"
              },
              "summary": "s16 elementwise multiplication with s16 output"
            },
            {
              "description": "s16 elementwise multiplication. The result of the multiplication is accumulated to the passed result buffer.\n\nSupported framework: TensorFlow Lite micro",
              "examples": [],
              "id": "arm_elementwise_mul_acc_s16",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_elementwise_mul_acc_s16",
              "params": [
                {
                  "description": "pointer to input vector 1",
                  "direction": "in",
                  "name": "input_1_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "pointer to input vector 2",
                  "direction": "in",
                  "name": "input_2_vect",
                  "type": "const int16_t *"
                },
                {
                  "description": "offset for input 1. Not used.",
                  "direction": "in",
                  "name": "input_1_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "offset for input 2. Not used.",
                  "direction": "in",
                  "name": "input_2_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "pointer to output vector",
                  "direction": "inout",
                  "name": "output",
                  "type": "int16_t *"
                },
                {
                  "description": "output offset. Not used.",
                  "direction": "in",
                  "name": "out_offset",
                  "type": "const int32_t"
                },
                {
                  "description": "output multiplier",
                  "direction": "in",
                  "name": "out_mult",
                  "type": "const int32_t"
                },
                {
                  "description": "output shift",
                  "direction": "in",
                  "name": "out_shift",
                  "type": "const int32_t"
                },
                {
                  "description": "minimum value to clamp output to. Min: -32768",
                  "direction": "in",
                  "name": "out_activation_min",
                  "type": "const int32_t"
                },
                {
                  "description": "maximum value to clamp output to. Max: 32767",
                  "direction": "in",
                  "name": "out_activation_max",
                  "type": "const int32_t"
                },
                {
                  "description": "number of samples",
                  "direction": "in",
                  "name": "block_size",
                  "type": "const int32_t"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns ARM_CMSIS_NN_SUCCESS"
                }
              ],
              "signature": "arm_cmsis_nn_status arm_elementwise_mul_acc_s16(\n    const int16_t *input_1_vect,\n    const int16_t *input_2_vect,\n    const int32_t input_1_offset,\n    const int32_t input_2_offset,\n    int16_t *output,\n    const int32_t out_offset,\n    const int32_t out_mult,\n    const int32_t out_shift,\n    const int32_t out_activation_min,\n    const int32_t out_activation_max,\n    const int32_t block_size\n)",
              "source": {
                "line": 3224,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3224"
              },
              "summary": "s16 elementwise multiplication."
            },
            {
              "description": "Check if a broadcast is required between 2 `cmsis_nn_dims`.\n\nCompares each dimension and returns 1 if any dimension does not match. This function does not check that broadcast rules are met.",
              "examples": [],
              "id": "arm_check_broadcast_required",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_check_broadcast_required",
              "params": [
                {
                  "description": "pointer to input tensor 1",
                  "direction": "in",
                  "name": "shape_1",
                  "type": "const cmsis_nn_dims *"
                },
                {
                  "description": "pointer to input tensor 2",
                  "direction": "in",
                  "name": "shape_2",
                  "type": "const cmsis_nn_dims *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "The function returns 1 if a broadcast is required, or 0 if not."
                }
              ],
              "signature": "static int32_t arm_check_broadcast_required(const cmsis_nn_dims *shape_1, const cmsis_nn_dims *shape_2)",
              "source": {
                "line": 3245,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3245"
              },
              "summary": "Check if a broadcast is required between 2 cmsisnndims."
            },
            {
              "description": "Reports whether the reduced axes of a 4-D tensor form one contiguous block followed by kept axes, as in a NHWC mean over H and W, and gives the flattened sizes. Axes of size 1 are ignored.",
              "examples": [],
              "id": "arm_reduce_get_middle_block_from_arrays",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_reduce_get_middle_block_from_arrays",
              "params": [
                {
                  "description": "4-element array {n, h, w, c}",
                  "direction": "in",
                  "name": "in_dims",
                  "type": "const int32_t"
                },
                {
                  "description": "4-element mask {axis_n, axis_h, axis_w, axis_c}",
                  "direction": "in",
                  "name": "axis_arr",
                  "type": "const int32_t"
                },
                {
                  "description": "Product of the dims before the reduced block",
                  "direction": "out",
                  "name": "outer",
                  "type": "int32_t *"
                },
                {
                  "description": "Product of the reduced dims",
                  "direction": "out",
                  "name": "reduce",
                  "type": "int32_t *"
                },
                {
                  "description": "Product of the dims after the reduced block",
                  "direction": "out",
                  "name": "inner",
                  "type": "int32_t *"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "1 if the input is [outer, reduce, inner] with the middle dim reduced and inner > 1, otherwise 0"
                }
              ],
              "signature": "static int32_t arm_reduce_get_middle_block_from_arrays(\n    const int32_t in_dims,\n    const int32_t axis_arr,\n    int32_t *outer,\n    int32_t *reduce,\n    int32_t *inner\n)",
              "source": {
                "line": 3310,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3310"
              },
              "summary": "Reports whether the reduced axes of a 4-D tensor form one contiguous block followed by kept axes, as in a NHWC mean over H and W, and gives the flattened sizes."
            },
            {
              "description": "",
              "examples": [],
              "id": "ARM_NN_SQRT_S16_TABLEFREE_SHIFT",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_SQRT_S16_TABLEFREE_SHIFT",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_SQRT_S16_TABLEFREE_SHIFT 14",
              "source": {
                "line": 3364,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3364"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "ARM_NN_SQRT_S16_TABLEFREE_MAGIC",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_SQRT_S16_TABLEFREE_MAGIC",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_SQRT_S16_TABLEFREE_MAGIC UINT32_C(0x5F5FB6C4)",
              "source": {
                "line": 3365,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3365"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "ARM_NN_SQRT_S16_TABLEFREE_K0",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_SQRT_S16_TABLEFREE_K0",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_SQRT_S16_TABLEFREE_K0 (-4.76426697f)",
              "source": {
                "line": 3366,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3366"
              },
              "summary": ""
            },
            {
              "description": "",
              "examples": [],
              "id": "ARM_NN_SQRT_S16_TABLEFREE_K1",
              "kind": "macro",
              "language": "c",
              "members": [],
              "name": "ARM_NN_SQRT_S16_TABLEFREE_K1",
              "params": [],
              "raises": [],
              "returns": [],
              "signature": "#define ARM_NN_SQRT_S16_TABLEFREE_K1 (-48.0000114f)",
              "source": {
                "line": 3367,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3367"
              },
              "summary": ""
            },
            {
              "description": "One element of `arm_sqrt_s16_tablefree()`: the float32 chain the MVE path evaluates per lane, so the two agree bit for bit on any IEEE-754 float32 implementation with round-to-nearest-even and a fused multiply-add (fmaf). Every product after the pre-scale either has two uses or feeds an fmaf or a conversion, never another lone multiply, so a compiler allowed to reassociate (-ffast-math) still has no chain to reorder, and no product feeds a bare add, so there is nothing to contract.",
              "examples": [],
              "id": "arm_nn_sqrt_s16_tablefree_element",
              "kind": "function",
              "language": "c",
              "members": [],
              "name": "arm_nn_sqrt_s16_tablefree_element",
              "params": [
                {
                  "description": "input code; values <= 0 give 0",
                  "direction": "in",
                  "name": "value",
                  "type": "const int32_t"
                },
                {
                  "description": "input_scale / (output_scale * output_scale) as float32",
                  "direction": "in",
                  "name": "scale",
                  "type": "const float"
                }
              ],
              "raises": [],
              "returns": [
                {
                  "description": "trunc(sqrt(value * scale)) saturated to 32767"
                }
              ],
              "signature": "static int16_t arm_nn_sqrt_s16_tablefree_element(const int32_t value, const float scale)",
              "source": {
                "line": 3381,
                "path": "Include/arm_nnsupportfunctions.h",
                "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnsupportfunctions.h#L3381"
              },
              "summary": "One element of armsqrts16tablefree(): the float32 chain the MVE path evaluates per lane, so the two agree bit for bit on any IEEE-754 float32 implementation wi…"
            }
          ]
        }
      ],
      "summary": "",
      "symbols": []
    }
  ],
  "name": "heliaCORE"
}
