Custom operators#

Overview#

A custom operator is a user-defined operator that extends the standard ONNX opset with functionality the framework does not provide natively. Models regularly contain operators that are:

  • Absent from the standard ONNX opset. Domain-specific math, novel attention variants, and geometric operations such as AffineGrid2D and GridSample2D are common examples.

  • Not yet supported by the NPU compiler.

  • Performance-critical fusions that you would rather map onto a hand-tuned AIE kernel than leave to default lowering.

Without a mechanism for custom operators, a model containing any of these either fails to compile or falls back to the CPU, which breaks end-to-end NPU acceleration.

The custom operator flow lets you supply your own AIE kernel, tiler, and schema, so that Vitis AI can compile, partition, and execute the operator on the NPU as it would any built-in one, keeping the whole graph on the device.

The following sections cover the artifacts a custom operator consists of, and how to register one with the compiler and the runtime. The pages linked at the end take it through Vitis AI compilation, ONNX Runtime inference, and on-board execution with x_plus_ml_vart.

Optionally, you can use the /vai-custom-op AI workflow, which automates this entire procedure; see Custom operator workflow.

Artifacts#

Every custom operator consists of three files:

Artifact

File pattern

Purpose

YAML configuration

customop_*.yaml

Central reference file for Vitis AI

AIE C++ kernel

customop_*.cpp

Per-core computation logic

Python tiler

customop_*_tiling.py

Memory partitioning and data movement

YAML configuration#

Vitis AI reads the YAML file when expanding a custom operator into AIE code. It identifies the tiler, and can carry per-core memory hints.

For example, customop_affinegrid2d/customop_affinegrid2d.yaml:

---
tiling: "customop_affinegrid2d_tiling.py"

Field

Purpose

tiling

Path to the Python tiler module

Note

Optional signature.inputs and signature.outputs entries can declare async flags, ONNX tensor layouts (onnx_tensor_layout), and DDR layouts (ddr_tensor_layout) for transposition and vectorization.

AIE kernel#

The C++ file implements the per-core AIE kernel, which Vitis AI invokes once per tile. It receives input and output buffer references, adf::input_buffer_conf and adf::output_buffer_conf, together with an lp_params integer array that the tiler populates.

A typical signature:

template <typename dtype_input, typename dtype_output>
__attribute__((noinline)) void affinegrid2d_kernel(
    adf::input_buffer_conf <dtype_input,  bpc_sync_0d> &__restrict in,
    adf::output_buffer_conf<dtype_output, bpc_sync_0d> &__restrict out,
    const int32_t (&lp_params)[<lp_size>]);

Important

The kernel symbol must match the name passed to set_kernel_function_name() in the tiling script.

For the constraints an AIE kernel must satisfy, see the architecture reference bundled with the release, described in Bundled examples and utilities.

Python tiler#

The tiler defines how the operator is partitioned across L3, L2, and L1 memory, and across the 4x4 AIE core grid. For the full tiling reference, see tiling.md in the bundled material described in Bundled examples and utilities.

Vitis AI calls a getTiling() entry point:

def getTiling(opInterface: tensor_expr.OperatorInfo,
              tiling: tensor_expr.AieConfig):
    stamp = tiling[0][0]  # Get the configurator for a stamp

    # 1. Declare buffers at each memory level (TensorVar.make ...)

    # 2. Tile tensors with .TileBy(index, size)

    # 3. Configure data movement
    #    stamp.set_l3_to_l2_transfer(...)
    #    stamp.set_l2_to_l1_transfer(...)
    #    stamp.set_l1_to_l2_transfer(...)
    #    stamp.set_l2_to_l3_transfer(...)

    # 4. Configure the kernel
    #    stamp.set_kernel_arguments([ifm_mk_var, wts_mk_var, ofm_mk_var])
    #    stamp.set_kernel_function_name("affinegrid2d_kernel")
    #    stamp.set_kernel_impl(Path("customop_kernel.cpp"))
    #    stamp.set_kernel_params([lp_params])
    #    stamp.set_kernel_nb_calls(num_ifm_tiles)

Registration#

Registration takes two independent steps, one for the compiler and one for the runtime. Both must agree on the operator’s domain and name.

Step

Where

Purpose

Compiler

vitisai_config.json

Registers the operator under the vaiml_partition pass and points to its YAML configuration, so the compiler knows which subgraphs to treat as custom operators.

Runtime

Python, before session creation

Declares the operator’s schema, that is its domain, name, and input and output counts, so that ONNX Runtime can parse the model graph.

Compiler#

Register each operator under the vaiml_partition pass. The key takes the form <domain>.<op_type_in_onnx>, and the value points to the YAML configuration:

{
  "passes": [
    {
      "name": "vaiml_partition",
      "plugin": "vaip-pass_vaiml_partition",
      "vaiml_config": {
        "device": "ve2",
        "custom_ops": {
          "mydomain.myAffineGrid2D": {
            "op_config": "customop_affinegrid2d/customop_affinegrid2d.yaml"
          },
          "mydomain.myScatterND": {
            "op_config": "customop_scatternd/customop_scatternd.yaml"
          },
          "mydomain.myGridSample2D": {
            "op_config": "customop_gridsample2d/customop_gridsample2d.yaml"
          }
        }
      }
    }
  ],
 "target": "VAIML",
 "targets": [
     {
         "name": "VAIML",
         "pass": [
             "init",
             "vaiml_partition"
         ]
     }
 ]
}

Runtime#

ONNX Runtime must know each operator’s schema before the inference session is created. Register schemas with register_dynamic_custom_ops_to_onnxruntime().

This example keeps the schema information in a helper file, custom_ops_config.json:

{
  "mydomain.myAffineGrid2D": {
    "domain": "mydomain",
    "name": "myAffineGrid2D",
    "inputs": 2,
    "outputs": 1,
    "op_config": "customop_affinegrid2d/customop_affinegrid2d.yaml"
  }
}

Then register from it:

import json
from onnxruntime_custom_ops import (
    register_dynamic_custom_ops_to_onnxruntime,
    vaiml_custom_op_schema,
)

with open("custom_ops_config.json") as f:
    custom_ops = json.load(f)

ops = [
    vaiml_custom_op_schema(
        domain=op["domain"],
        name=op["name"],
        nb_inputs=op["inputs"],
        nb_outputs=op["outputs"],
    )
    for op in custom_ops.values()
]
register_dynamic_custom_ops_to_onnxruntime(ops)

Important

The registered name must match the operator type used inside the ONNX model exactly. A mismatch produces:

Fatal error: mydomain:<opname>(-1) is not a registered function/op.

Once the schemas are registered, create the inference session with VitisAIExecutionProvider, setting config_file to vitisai_config.json. The compiler partitions the graph, compiles each custom operator subgraph according to its YAML configuration, and writes the compiled artifacts to <cache_dir>/<cache_key>/.

Bundled examples and utilities#

Tutorials, worked examples, and supporting utilities ship in the release container, under:

/usr/local/lib/python3.12/dist-packages/flexml/flexml_extras/ai_utils

Location under ai_utils

Contents

skills/vai-custom-op-implementation/tutorial/

Complete worked custom operator examples, including mul, negate, topk, and softmax. Useful as starting points for manual operator development.

skills/vai-custom-op-implementation/scripts/

Python utilities for compiling and simulating an operator, and for cutting and stitching ONNX graphs. Runnable directly from the command line.

skills/vai-custom-op-implementation/README.md

AIE architecture reference covering the constraints a custom operator must satisfy.

skills/vai-custom-op-implementation/tiling.md

Tiling reference covering memory partitioning and data movement.