{
  "$schema": "https://ambiqai.github.io/helia-ui/schema/reference-model-1.json",
  "generatedFrom": {
    "sourceCommit": "5f3fed9f21a57390cc7f00f77a37db8f5f110cb8",
    "tool": "doxyref",
    "version": "1.17.0"
  },
  "language": "c",
  "modules": [
    {
      "description": "Collection of convolution, depthwise convolution functions and their variants.\n\nThe convolution is implemented in 2 steps: im2col and General Matrix Multiplication(GEMM)\n\nim2col is a process of converting each patch of image data into a column. After im2col, the convolution is computed as matrix-matrix multiplication.\n\nTo reduce the memory footprint, the im2col is performed partially. Each iteration, only a few column (i.e., patches) are generated followed by GEMM.",
      "name": "Convolution Functions",
      "path": "heliaCORE.NNConv",
      "submodules": [],
      "summary": "Collection of convolution, depthwise convolution functions and their variants.",
      "symbols": [
        {
          "description": "Depthwise convolution, NHWC layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::",
          "examples": [],
          "id": "arm_depthwise_nhwc_conv_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_nhwc_conv_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions in NHWC format.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions in NHWC-compatible depthwise format.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "kernel",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions in NHWC format.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_depthwise_nhwc_conv_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output\n)",
          "source": {
            "line": 79,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L79"
          },
          "summary": "Depthwise convolution, NHWC layout."
        },
        {
          "description": "Depthwise convolution, dispatch by layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::",
          "examples": [],
          "id": "arm_depthwise_conv_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "kernel",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_depthwise_conv_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 119,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L119"
          },
          "summary": "Depthwise convolution, dispatch by layout."
        },
        {
          "description": "Depthwise convolution wrapper using the CMSIS-NN baseline path.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::",
          "examples": [],
          "id": "arm_depthwise_conv_wrapper_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_wrapper_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Depthwise convolution parameters.",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "kernel",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output",
              "type": "float32_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_depthwise_conv_wrapper_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output\n)",
          "source": {
            "line": 158,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L158"
          },
          "summary": "Depthwise convolution wrapper using the CMSIS-NN baseline path."
        },
        {
          "description": "Get the temporary buffer size required by depthwise convolution.\n\n:::note\nOnly one route reads scratch: on MVE builds, an NHWC depthwise with ch_mult != 1, a single input channel and at least CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD output channels can run as a regular convolution, which needs the repacked filter, `ROUND_UP(output_dims->c, 4) * filter_dims->h * filter_dims->w * sizeof(float32_t)` bytes (`ROUND_UP(output_dims->c, 8)` and `sizeof(float16_t)` for `_f16`), plus `arm_convolve_wrapper_f32_get_buffer_size` (`_f16`) for that convolution. The query reserves that for every such layer; one an exact-shape specialization takes first runs without it, so the size is an upper bound there. Every other layer  the `ch_mult == 1` direct kernel and the generic kernel  runs without scratch and the query returns 0 (AmbiqAI/ns-cmsis-nn#448, #625).\n\n:::",
          "examples": [],
          "id": "arm_depthwise_conv_f32_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_f32_get_buffer_size",
          "params": [
            {
              "description": "Depthwise convolution parameters.",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Tensor layout selector.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_depthwise_conv_f32_get_buffer_size(\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 189,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L189"
          },
          "summary": "Get the temporary buffer size required by depthwise convolution."
        },
        {
          "description": "Get the buffer size required by the depthwise convolution wrapper.",
          "examples": [],
          "id": "arm_depthwise_conv_wrapper_f32_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_wrapper_f32_get_buffer_size",
          "params": [
            {
              "description": "Depthwise convolution parameters.",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_depthwise_conv_wrapper_f32_get_buffer_size(\n    const cmsis_nn_dw_conv_params_f32 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
          "source": {
            "line": 205,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L205"
          },
          "summary": "Get the buffer size required by the depthwise convolution wrapper."
        },
        {
          "description": "Convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
          "examples": [],
          "id": "arm_convolve_nhwc_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_nhwc_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_nhwc_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
          "source": {
            "line": 232,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L232"
          },
          "summary": "Convolution, NHWC layout."
        },
        {
          "description": "Convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
          "examples": [],
          "id": "arm_convolve_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 266,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L266"
          },
          "summary": "Convolution, dispatch by layout."
        },
        {
          "description": "Convolution wrapper using the CMSIS-NN baseline path.",
          "examples": [],
          "id": "arm_convolve_wrapper_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_wrapper_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_wrapper_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
          "source": {
            "line": 294,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L294"
          },
          "summary": "Convolution wrapper using the CMSIS-NN baseline path."
        },
        {
          "description": "1x1 convolution, NHWC layout.",
          "examples": [],
          "id": "arm_convolve_1x1_nhwc_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1x1_nhwc_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1x1_nhwc_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
          "source": {
            "line": 321,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L321"
          },
          "summary": "1x1 convolution, NHWC layout."
        },
        {
          "description": "1x1 convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1x1_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1x1_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1x1_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 354,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L354"
          },
          "summary": "1x1 convolution, dispatch by layout."
        },
        {
          "description": "1xN convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1_x_n_nhwc_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1_x_n_nhwc_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
          "source": {
            "line": 391,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L391"
          },
          "summary": "1xN convolution, NHWC layout."
        },
        {
          "description": "1xN convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1_x_n_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1_x_n_f32",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 428,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L428"
          },
          "summary": "1xN convolution, dispatch by layout."
        },
        {
          "description": "Get the temporary buffer size required by convolution.\n\n:::note\nWhen `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still reports only the temporary input/im2col scratch requirement. Any offline-packed filter storage is expected to be provided by the caller.\n\n:::",
          "examples": [],
          "id": "arm_convolve_f32_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_f32_get_buffer_size",
          "params": [
            {
              "description": "Convolution parameters.",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Tensor layout selector.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_convolve_f32_get_buffer_size(\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 456,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L456"
          },
          "summary": "Get the temporary buffer size required by convolution."
        },
        {
          "description": "Get the buffer size required by the convolution wrapper.",
          "examples": [],
          "id": "arm_convolve_wrapper_f32_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_wrapper_f32_get_buffer_size",
          "params": [
            {
              "description": "Convolution parameters.",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_convolve_wrapper_f32_get_buffer_size(\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
          "source": {
            "line": 472,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L472"
          },
          "summary": "Get the buffer size required by the convolution wrapper."
        },
        {
          "description": "Get the buffer size required by 1x1 convolution.\n\n:::note\nReturns `0` for the unity-stride no-pack path. For non-unity-stride NHWC 1x1 convolution, the returned scratch size enables the packed-tile + GEMM path. When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still excludes the offline-packed filter storage itself.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1x1_f32_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1x1_f32_get_buffer_size",
          "params": [
            {
              "description": "Convolution parameters.",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Tensor layout selector.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_convolve_1x1_f32_get_buffer_size(\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 493,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L493"
          },
          "summary": "Get the buffer size required by 1x1 convolution."
        },
        {
          "description": "Get the buffer size required by 1xN convolution.",
          "examples": [],
          "id": "arm_convolve_1_x_n_f32_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1_x_n_f32_get_buffer_size",
          "params": [
            {
              "description": "Convolution parameters.",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Tensor layout selector.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_convolve_1_x_n_f32_get_buffer_size(\n    const cmsis_nn_conv_params_f32 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 510,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L510"
          },
          "summary": "Get the buffer size required by 1xN convolution."
        },
        {
          "description": "Transpose convolution wrapper using the CMSIS-NN baseline path.",
          "examples": [],
          "id": "arm_transpose_conv_wrapper_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_wrapper_f32",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Output context. Unused; may be NULL.",
              "direction": "in",
              "name": "output_ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions.",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_transpose_conv_wrapper_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 1540,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1540"
          },
          "summary": "Transpose convolution wrapper using the CMSIS-NN baseline path."
        },
        {
          "description": "Transpose convolution, NHWC layout.",
          "examples": [],
          "id": "arm_transpose_conv_nhwc_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_nhwc_f32",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Output context. Unused; may be NULL.",
              "direction": "in",
              "name": "output_ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions.",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_transpose_conv_nhwc_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data\n)",
          "source": {
            "line": 1570,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1570"
          },
          "summary": "Transpose convolution, NHWC layout."
        },
        {
          "description": "Transpose convolution, dispatch by layout.",
          "examples": [],
          "id": "arm_transpose_conv_f32",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_f32",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Output context. Unused; may be NULL.",
              "direction": "in",
              "name": "output_ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float32_t *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float32_t *"
            },
            {
              "description": "Bias tensor dimensions.",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float32_t *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float32_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_transpose_conv_f32(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float32_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float32_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float32_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float32_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 1600,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1600"
          },
          "summary": "Transpose convolution, dispatch by layout."
        },
        {
          "description": "Get the temporary buffer size required by transpose convolution.",
          "examples": [],
          "id": "arm_transpose_conv_f32_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_f32_get_buffer_size",
          "params": [
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "out_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_transpose_conv_f32_get_buffer_size(\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *out_dims\n)",
          "source": {
            "line": 1623,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1623"
          },
          "summary": "Get the temporary buffer size required by transpose convolution."
        },
        {
          "description": "Get the reverse-convolution workspace size used by transpose convolution helpers.",
          "examples": [],
          "id": "arm_transpose_conv_f32_get_reverse_conv_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_f32_get_reverse_conv_buffer_size",
          "params": [
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f32 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no reverse-convolution buffer is needed."
            }
          ],
          "signature": "int32_t arm_transpose_conv_f32_get_reverse_conv_buffer_size(\n    const cmsis_nn_transpose_conv_params_f32 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims\n)",
          "source": {
            "line": 1638,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L1638"
          },
          "summary": "Get the reverse-convolution workspace size used by transpose convolution helpers."
        },
        {
          "description": "Depthwise convolution, NHWC layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::",
          "examples": [],
          "id": "arm_depthwise_nhwc_conv_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_nhwc_conv_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions in NHWC format.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions in NHWC-compatible depthwise format.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "kernel",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions in NHWC format.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_depthwise_nhwc_conv_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
          "source": {
            "line": 2224,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2224"
          },
          "summary": "Depthwise convolution, NHWC layout."
        },
        {
          "description": "Depthwise convolution, NHWC layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_nhwc_conv_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_nhwc_conv_f16.\n\n:::",
          "examples": [],
          "id": "arm_depthwise_nhwc_conv_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_nhwc_conv_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions in NHWC format.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions in NHWC-compatible depthwise format.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "kernel",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions in NHWC format.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_depthwise_nhwc_conv_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
          "source": {
            "line": 2243,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2243"
          },
          "summary": "Depthwise convolution, NHWC layout."
        },
        {
          "description": "Depthwise convolution, dispatch by layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::",
          "examples": [],
          "id": "arm_depthwise_conv_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "kernel",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_depthwise_conv_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2273,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2273"
          },
          "summary": "Depthwise convolution, dispatch by layout."
        },
        {
          "description": "Depthwise convolution, dispatch by layout.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_conv_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_conv_f16.\n\n:::",
          "examples": [],
          "id": "arm_depthwise_conv_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Depthwise convolution parameters (stride, padding, dilation, channel multiplier and activation clamp).",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "kernel",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_depthwise_conv_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2293,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2293"
          },
          "summary": "Depthwise convolution, dispatch by layout."
        },
        {
          "description": "Depthwise convolution wrapper using the CMSIS-NN baseline path.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::",
          "examples": [],
          "id": "arm_depthwise_conv_wrapper_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_wrapper_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Depthwise convolution parameters.",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "kernel",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_depthwise_conv_wrapper_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
          "source": {
            "line": 2324,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2324"
          },
          "summary": "Depthwise convolution wrapper using the CMSIS-NN baseline path."
        },
        {
          "description": "Depthwise convolution wrapper using the CMSIS-NN baseline path.\n\n:::note\nWhen `ctx->buf` is used for internal kernel repacking, it must be aligned to the element type stored in scratch: at least 4-byte aligned for `float32_t` and at least 2-byte aligned for float16_t.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). Every route accumulates the bias and every tap in float32 on every leg. A NaN input tap or weight does not propagate: the `ch_mult == 1` direct kernel's MVE leg clamps it to the activation minimum (`vmaxnm` / `vminnm`); the MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) clamps it to the activation minimum (`arm_nn_clamp_mve_f32`); its scalar leg and the `ch_mult > 1` generic kernel clamp it to the activation maximum (`ARM_NN_CLAMP`). That is the pre-#448 behavior of these routes; unifying it with the float16 scalar leg under the #334 promise is a separate issue.\n\n:::\n\n:::note\nAccumulation and NaN, per leg (AmbiqAI/ns-cmsis-nn#448). MVE leg: the `ch_mult == 1` direct kernel (lanes are channels, taps row by row) uses blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. It clamps a NaN to the activation minimum (`vmaxnm` / `vminnm`). Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct kernel accumulates in float32 and rounds to float16 once at the store (#449), and a NaN propagates through `arm_nn_clamp_scalar_f16`  unlike the float32 scalar leg, which clamps it to a bound. The MVE to-convolution route (input channels 1, output channels 8 or more, ctx supplied) goes through arm_nn_mat_mult_nt_n_packed_f16, with the same blockwise rule over every tap, padded ones included, and clamps a NaN to the activation minimum (`arm_nn_clamp_mve_f16`). The `ch_mult > 1` generic kernel accumulates in float16 with the same blockwise rule on MVE builds; on the scalar legs it accumulates in float32, bias included, and rounds to float16 once (#645), the same on both entries. It clamps a NaN to the activation maximum (`arm_nn_clamp_f16h`) on every leg. Unifying these under the #334 promise is a separate issue.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_depthwise_conv_wrapper_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_depthwise_conv_wrapper_f16.\n\n:::",
          "examples": [],
          "id": "arm_depthwise_conv_wrapper_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_wrapper_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Depthwise convolution parameters.",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "kernel",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_depthwise_conv_wrapper_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *kernel,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output\n)",
          "source": {
            "line": 2343,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2343"
          },
          "summary": "Depthwise convolution wrapper using the CMSIS-NN baseline path."
        },
        {
          "description": "Get the temporary buffer size required by depthwise convolution.\n\n:::note\nOnly one route reads scratch: on MVE builds, an NHWC depthwise with ch_mult != 1, a single input channel and at least CONVERT_DW_CONV_WITH_ONE_INPUT_CH_AND_OUTPUT_CH_ABOVE_THRESHOLD output channels can run as a regular convolution, which needs the repacked filter, `ROUND_UP(output_dims->c, 4) * filter_dims->h * filter_dims->w * sizeof(float32_t)` bytes (`ROUND_UP(output_dims->c, 8)` and `sizeof(float16_t)` for `_f16`), plus `arm_convolve_wrapper_f32_get_buffer_size` (`_f16`) for that convolution. The query reserves that for every such layer; one an exact-shape specialization takes first runs without it, so the size is an upper bound there. Every other layer  the `ch_mult == 1` direct kernel and the generic kernel  runs without scratch and the query returns 0 (AmbiqAI/ns-cmsis-nn#448, #625).\n\n:::",
          "examples": [],
          "id": "arm_depthwise_conv_f16_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_f16_get_buffer_size",
          "params": [
            {
              "description": "Depthwise convolution parameters.",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Tensor layout selector.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_depthwise_conv_f16_get_buffer_size(\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2357,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2357"
          },
          "summary": "Get the temporary buffer size required by depthwise convolution."
        },
        {
          "description": "Get the buffer size required by the depthwise convolution wrapper.",
          "examples": [],
          "id": "arm_depthwise_conv_wrapper_f16_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_depthwise_conv_wrapper_f16_get_buffer_size",
          "params": [
            {
              "description": "Depthwise convolution parameters.",
              "direction": "in",
              "name": "dw_conv_params",
              "type": "const cmsis_nn_dw_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_depthwise_conv_wrapper_f16_get_buffer_size(\n    const cmsis_nn_dw_conv_params_f16 *dw_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
          "source": {
            "line": 2366,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2366"
          },
          "summary": "Get the buffer size required by the depthwise convolution wrapper."
        },
        {
          "description": "Convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
          "examples": [],
          "id": "arm_convolve_nhwc_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_nhwc_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_nhwc_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
          "source": {
            "line": 2374,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2374"
          },
          "summary": "Convolution, NHWC layout."
        },
        {
          "description": "Convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_nhwc_f16.\n\n:::",
          "examples": [],
          "id": "arm_convolve_nhwc_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_nhwc_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_nhwc_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
          "source": {
            "line": 2393,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2393"
          },
          "summary": "Convolution, NHWC layout."
        },
        {
          "description": "Convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct OHWI / NT_N_PACKED fallback accumulates bias and every tap in float32 and rounds to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #457); the 1x1, 1xN and patch-GEMM paths go through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16, whose scalar legs do the same, as do the 1xN no-padding OHWI region and the k=3 / k=5 conv1d specializations (#465). MVE leg: the direct small-C kernel accumulates in float32 (widened lanes); the direct OHWI / NT_N_PACKED fallback, every matmul-backed path (1x1, 1xN, patch-GEMM), the 1xN no-padding region and the conv1d specializations use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. Where a dot product spreads its taps over the lanes of one vector (OHWI rows, the contiguous-K matmul, the conv1d k=3 / k=5 OHWI kernels), each lane's own taps form its blocks and, once the reduction exceeds 32 taps, each block's lanes are widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (the first block sets it); the four pair accumulators are summed once as (0+1) + (2+3), the bias is added in float32 and the total rounds once. The k=3 / k=5 kernels close a block on a whole input-channel step (30 taps per lane). An output's taps are the ones its kernel multiplies: the direct fallback skips padded taps, so an edge output counts only its in-range taps, while patch-GEMM and the 1xN padded regions multiply a zero-padded patch and count its padded taps too. Patch-GEMM runs only when ctx provides its scratch, so an edge output's value can depend on whether ctx->buf is given. The fold's order is fixed; the float16 reduction of a dot of at most 32 taps is left to the compiler, which may reorder it under -ffast-math, as before #586.\n\n:::",
          "examples": [],
          "id": "arm_convolve_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2431,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2431"
          },
          "summary": "Convolution, dispatch by layout."
        },
        {
          "description": "Convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, every convolution path, including the 1xN kernels and the generic fallback that runs without scratch, interprets `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): the direct OHWI / NT_N_PACKED fallback accumulates bias and every tap in float32 and rounds to float16 once at the store (AmbiqAI/ns-cmsis-nn#449, #457); the 1x1, 1xN and patch-GEMM paths go through arm_nn_mat_mult_nt_t_f16 / arm_nn_mat_mult_nt_n_packed_f16, whose scalar legs do the same, as do the 1xN no-padding OHWI region and the k=3 / k=5 conv1d specializations (#465). MVE leg: the direct small-C kernel accumulates in float32 (widened lanes); the direct OHWI / NT_N_PACKED fallback, every matmul-backed path (1x1, 1xN, patch-GEMM), the 1xN no-padding region and the conv1d specializations use blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. Where a dot product spreads its taps over the lanes of one vector (OHWI rows, the contiguous-K matmul, the conv1d k=3 / k=5 OHWI kernels), each lane's own taps form its blocks and, once the reduction exceeds 32 taps, each block's lanes are widened and lanes 2j and 2j+1 added in float32 into pair accumulator j (the first block sets it); the four pair accumulators are summed once as (0+1) + (2+3), the bias is added in float32 and the total rounds once. The k=3 / k=5 kernels close a block on a whole input-channel step (30 taps per lane). An output's taps are the ones its kernel multiplies: the direct fallback skips padded taps, so an edge output counts only its in-range taps, while patch-GEMM and the 1xN padded regions multiply a zero-padded patch and count its padded taps too. Patch-GEMM runs only when ctx provides its scratch, so an edge output's value can depend on whether ctx->buf is given. The fold's order is fixed; the float16 reduction of a dot of at most 32 taps is left to the compiler, which may reorder it under -ffast-math, as before #586.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_f16.\n\n:::",
          "examples": [],
          "id": "arm_convolve_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2451,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2451"
          },
          "summary": "Convolution, dispatch by layout."
        },
        {
          "description": "Convolution wrapper using the CMSIS-NN baseline path.",
          "examples": [],
          "id": "arm_convolve_wrapper_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_wrapper_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_wrapper_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
          "source": {
            "line": 2466,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2466"
          },
          "summary": "Convolution wrapper using the CMSIS-NN baseline path."
        },
        {
          "description": "Convolution wrapper using the CMSIS-NN baseline path.\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_wrapper_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_wrapper_f16.\n\n:::",
          "examples": [],
          "id": "arm_convolve_wrapper_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_wrapper_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_wrapper_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
          "source": {
            "line": 2485,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2485"
          },
          "summary": "Convolution wrapper using the CMSIS-NN baseline path."
        },
        {
          "description": "1x1 convolution, NHWC layout.",
          "examples": [],
          "id": "arm_convolve_1x1_nhwc_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1x1_nhwc_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1x1_nhwc_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
          "source": {
            "line": 2499,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2499"
          },
          "summary": "1x1 convolution, NHWC layout."
        },
        {
          "description": "1x1 convolution, NHWC layout.\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1x1_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1x1_nhwc_f16.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1x1_nhwc_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1x1_nhwc_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1x1_nhwc_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
          "source": {
            "line": 2518,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2518"
          },
          "summary": "1x1 convolution, NHWC layout."
        },
        {
          "description": "1x1 convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1x1_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1x1_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1x1_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2532,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2532"
          },
          "summary": "1x1 convolution, dispatch by layout."
        },
        {
          "description": "1x1 convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, the matmul-backed 1x1 convolution paths interpret `filter_data` as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1x1_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1x1_f16.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1x1_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1x1_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1x1_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2552,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2552"
          },
          "summary": "1x1 convolution, dispatch by layout."
        },
        {
          "description": "1xN convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1_x_n_nhwc_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1_x_n_nhwc_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
          "source": {
            "line": 2567,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2567"
          },
          "summary": "1xN convolution, NHWC layout."
        },
        {
          "description": "1xN convolution, NHWC layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1_x_n_nhwc_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1_x_n_nhwc_f16.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1_x_n_nhwc_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1_x_n_nhwc_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format: [N, H, W, C_IN].",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format: [C_OUT, HK, WK, C_IN].",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format: [N, H, W, C_OUT].",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_nhwc_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
          "source": {
            "line": 2586,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2586"
          },
          "summary": "1xN convolution, NHWC layout."
        },
        {
          "description": "1xN convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and every product accumulate in float32 and round to float16 once (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: the padded regions go through the matmul helpers and the no-padding region through a strided kernel, all with blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. The padded regions multiply a zero-padded patch row, so their outputs count the padded taps as well.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1_x_n_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1_x_n_f16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2610,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2610"
          },
          "summary": "1xN convolution, dispatch by layout."
        },
        {
          "description": "1xN convolution, dispatch by layout.\n\n:::note\nWhen `conv_params->weight_format` is set to `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, `filter_data` is interpreted as an already prepacked `NTxN` RHS buffer instead of the standard public filter layout. Every output position, including the no-pad middle region that OHWI filters feed to the strided direct kernel, is then packed into scratch and multiplied by the format-aware matmul, so a packed 1xN layer with little or no padding runs slower than its OHWI equivalent.\n\n:::\n\n:::note\nAccumulation width per leg. Scalar leg (non-MVE builds and ARM_MATH_AUTOVECTORIZE): bias and every product accumulate in float32 and round to float16 once (AmbiqAI/ns-cmsis-nn#449, #465). MVE leg: the padded regions go through the matmul helpers and the no-padding region through a strided kernel, all with blockwise float16 accumulation (AmbiqAI/ns-cmsis-nn#586, superseding #446's float16-lane choice for the MVE legs): in the kernel's own tap order an accumulator lane sums at most 32 taps in float16 (the bias, where the kernel starts from it, opens the first block), then the partial is widened exactly and added into a float32 accumulator; the float32 sum rounds to float16 once, before the clamp. An accumulator of at most 32 taps gives exactly the float16-lane result. The `_acc16` entry keeps float16 lanes throughout. The padded regions multiply a zero-padded patch row, so their outputs count the padded taps as well.\n\n:::\n\n:::note\nFloat16-lane entry (AmbiqAI/ns-cmsis-nn#586): the MVE legs run with no blockwise fold, exactly as arm_convolve_1_x_n_f16 did before #586 (float16 accumulator lanes wherever it used them), for callers that trade accuracy on long reductions for speed. Same arguments, scratch buffer (and sizer), return codes and scalar leg as arm_convolve_1_x_n_f16.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1_x_n_f16_acc16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1_x_n_f16_acc16",
          "params": [
            {
              "description": "Function context that may hold a temporary scratch buffer.",
              "direction": "inout",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Convolution parameters (stride, padding, dilation and activation clamp).",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions. Format: [C_OUT].",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions. Format depends on `layout`.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_convolve_1_x_n_f16_acc16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2630,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2630"
          },
          "summary": "1xN convolution, dispatch by layout."
        },
        {
          "description": "Get the temporary buffer size required by convolution.\n\n:::note\nWhen `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still reports only the temporary input/im2col scratch requirement. Any offline-packed filter storage is expected to be provided by the caller.\n\n:::",
          "examples": [],
          "id": "arm_convolve_f16_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_f16_get_buffer_size",
          "params": [
            {
              "description": "Convolution parameters.",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Tensor layout selector.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_convolve_f16_get_buffer_size(\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2645,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2645"
          },
          "summary": "Get the temporary buffer size required by convolution."
        },
        {
          "description": "Get the buffer size required by the convolution wrapper.",
          "examples": [],
          "id": "arm_convolve_wrapper_f16_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_wrapper_f16_get_buffer_size",
          "params": [
            {
              "description": "Convolution parameters.",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_convolve_wrapper_f16_get_buffer_size(\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims\n)",
          "source": {
            "line": 2654,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2654"
          },
          "summary": "Get the buffer size required by the convolution wrapper."
        },
        {
          "description": "Get the buffer size required by 1x1 convolution.\n\n:::note\nReturns `0` for the unity-stride no-pack path. For non-unity-stride NHWC 1x1 convolution, the returned scratch size enables the packed-tile + GEMM path. When `conv_params->weight_format` is `ARM_NN_WEIGHT_FORMAT_NT_N_PACKED`, this still excludes the offline-packed filter storage itself.\n\n:::",
          "examples": [],
          "id": "arm_convolve_1x1_f16_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1x1_f16_get_buffer_size",
          "params": [
            {
              "description": "Convolution parameters.",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Tensor layout selector.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_convolve_1x1_f16_get_buffer_size(\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2662,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2662"
          },
          "summary": "Get the buffer size required by 1x1 convolution."
        },
        {
          "description": "Get the buffer size required by 1xN convolution.",
          "examples": [],
          "id": "arm_convolve_1_x_n_f16_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_convolve_1_x_n_f16_get_buffer_size",
          "params": [
            {
              "description": "Convolution parameters.",
              "direction": "in",
              "name": "conv_params",
              "type": "const cmsis_nn_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Tensor layout selector.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_convolve_1_x_n_f16_get_buffer_size(\n    const cmsis_nn_conv_params_f16 *conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *output_dims,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 2671,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L2671"
          },
          "summary": "Get the buffer size required by 1xN convolution."
        },
        {
          "description": "Transpose convolution wrapper using the CMSIS-NN baseline path.",
          "examples": [],
          "id": "arm_transpose_conv_wrapper_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_wrapper_f16",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Output context. Unused; may be NULL.",
              "direction": "in",
              "name": "output_ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions.",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_transpose_conv_wrapper_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 3354,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3354"
          },
          "summary": "Transpose convolution wrapper using the CMSIS-NN baseline path."
        },
        {
          "description": "Transpose convolution, NHWC layout.",
          "examples": [],
          "id": "arm_transpose_conv_nhwc_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_nhwc_f16",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Output context. Unused; may be NULL.",
              "direction": "in",
              "name": "output_ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions.",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_transpose_conv_nhwc_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data\n)",
          "source": {
            "line": 3370,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3370"
          },
          "summary": "Transpose convolution, NHWC layout."
        },
        {
          "description": "Transpose convolution, dispatch by layout.",
          "examples": [],
          "id": "arm_transpose_conv_f16",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_f16",
          "params": [
            {
              "description": "Function context. Unused; may be NULL.",
              "direction": "in",
              "name": "ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Output context. Unused; may be NULL.",
              "direction": "in",
              "name": "output_ctx",
              "type": "const cmsis_nn_context *"
            },
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the input tensor data.",
              "direction": "in",
              "name": "input_data",
              "type": "const float16_t *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the filter tensor data.",
              "direction": "in",
              "name": "filter_data",
              "type": "const float16_t *"
            },
            {
              "description": "Bias tensor dimensions.",
              "direction": "in",
              "name": "bias_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Optional bias tensor data.",
              "direction": "in",
              "name": "bias_data",
              "type": "const float16_t *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "output_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Pointer to the output tensor data.",
              "direction": "out",
              "name": "output_data",
              "type": "float16_t *"
            },
            {
              "description": "Tensor layout selector. Current float APIs require `ARM_NN_LAYOUT_NHWC`.",
              "direction": "in",
              "name": "layout",
              "type": "arm_nn_tensor_layout"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "`ARM_CMSIS_NN_SUCCESS` on success or `ARM_CMSIS_NN_ARG_ERROR` on invalid arguments."
            }
          ],
          "signature": "arm_cmsis_nn_status arm_transpose_conv_f16(\n    const cmsis_nn_context *ctx,\n    const cmsis_nn_context *output_ctx,\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const float16_t *input_data,\n    const cmsis_nn_dims *filter_dims,\n    const float16_t *filter_data,\n    const cmsis_nn_dims *bias_dims,\n    const float16_t *bias_data,\n    const cmsis_nn_dims *output_dims,\n    float16_t *output_data,\n    arm_nn_tensor_layout layout\n)",
          "source": {
            "line": 3385,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3385"
          },
          "summary": "Transpose convolution, dispatch by layout."
        },
        {
          "description": "Get the temporary buffer size required by transpose convolution.",
          "examples": [],
          "id": "arm_transpose_conv_f16_get_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_f16_get_buffer_size",
          "params": [
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Output tensor dimensions.",
              "direction": "in",
              "name": "out_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no scratch buffer is needed."
            }
          ],
          "signature": "int32_t arm_transpose_conv_f16_get_buffer_size(\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims,\n    const cmsis_nn_dims *out_dims\n)",
          "source": {
            "line": 3401,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3401"
          },
          "summary": "Get the temporary buffer size required by transpose convolution."
        },
        {
          "description": "Get the reverse-convolution workspace size used by transpose convolution helpers.",
          "examples": [],
          "id": "arm_transpose_conv_f16_get_reverse_conv_buffer_size",
          "kind": "function",
          "language": "c",
          "members": [],
          "name": "arm_transpose_conv_f16_get_reverse_conv_buffer_size",
          "params": [
            {
              "description": "Transpose convolution parameters.",
              "direction": "in",
              "name": "transpose_conv_params",
              "type": "const cmsis_nn_transpose_conv_params_f16 *"
            },
            {
              "description": "Input tensor dimensions.",
              "direction": "in",
              "name": "input_dims",
              "type": "const cmsis_nn_dims *"
            },
            {
              "description": "Filter tensor dimensions.",
              "direction": "in",
              "name": "filter_dims",
              "type": "const cmsis_nn_dims *"
            }
          ],
          "raises": [],
          "returns": [
            {
              "description": "Required buffer size in bytes, or 0 when no reverse-convolution buffer is needed."
            }
          ],
          "signature": "int32_t arm_transpose_conv_f16_get_reverse_conv_buffer_size(\n    const cmsis_nn_transpose_conv_params_f16 *transpose_conv_params,\n    const cmsis_nn_dims *input_dims,\n    const cmsis_nn_dims *filter_dims\n)",
          "source": {
            "line": 3410,
            "path": "Include/arm_nnfunctions_flt.h",
            "url": "https://github.com/AmbiqAI/ns-cmsis-nn/blob/5f3fed9f21a57390cc7f00f77a37db8f5f110cb8/Include/arm_nnfunctions_flt.h#L3410"
          },
          "summary": "Get the reverse-convolution workspace size used by transpose convolution helpers."
        }
      ]
    }
  ],
  "name": "heliaCORE"
}
