Custom operators#
Overview#
A custom operator is a user-defined operator that extends the standard ONNX opset with functionality the framework does not provide natively. Models regularly contain operators that are:
Absent from the standard ONNX opset. Domain-specific math, novel attention variants, and geometric operations such as
AffineGrid2DandGridSample2Dare common examples.Not yet supported by the NPU compiler.
Performance-critical fusions that you would rather map onto a hand-tuned AIE kernel than leave to default lowering.
Without a mechanism for custom operators, a model containing any of these either fails to compile or falls back to the CPU, which breaks end-to-end NPU acceleration.
The custom operator flow lets you supply your own AIE kernel, tiler, and schema, so that Vitis AI can compile, partition, and execute the operator on the NPU as it would any built-in one, keeping the whole graph on the device.
The following sections cover the artifacts a custom operator consists of, and how to
register one with the compiler and the runtime. The pages linked at the end take
it through Vitis AI compilation, ONNX Runtime inference, and on-board execution
with x_plus_ml_vart.
Optionally, you can use the /vai-custom-op AI workflow, which automates this entire procedure; see
Custom operator workflow.
Artifacts#
Every custom operator consists of three files:
Artifact |
File pattern |
Purpose |
|---|---|---|
YAML configuration |
|
Central reference file for Vitis AI |
AIE C++ kernel |
|
Per-core computation logic |
Python tiler |
|
Memory partitioning and data movement |
YAML configuration#
Vitis AI reads the YAML file when expanding a custom operator into AIE code. It identifies the tiler, and can carry per-core memory hints.
For example, customop_affinegrid2d/customop_affinegrid2d.yaml:
---
tiling: "customop_affinegrid2d_tiling.py"
Field |
Purpose |
|---|---|
|
Path to the Python tiler module |
Note
Optional signature.inputs and signature.outputs entries can declare
async flags, ONNX tensor layouts (onnx_tensor_layout), and DDR layouts
(ddr_tensor_layout) for transposition and vectorization.
AIE kernel#
The C++ file implements the per-core AIE kernel, which Vitis AI invokes once per
tile. It receives input and output buffer references,
adf::input_buffer_conf and adf::output_buffer_conf, together with an
lp_params integer array that the tiler populates.
A typical signature:
template <typename dtype_input, typename dtype_output>
__attribute__((noinline)) void affinegrid2d_kernel(
adf::input_buffer_conf <dtype_input, bpc_sync_0d> &__restrict in,
adf::output_buffer_conf<dtype_output, bpc_sync_0d> &__restrict out,
const int32_t (&lp_params)[<lp_size>]);
Important
The kernel symbol must match the name passed to
set_kernel_function_name() in the tiling script.
For the constraints an AIE kernel must satisfy, see the architecture reference bundled with the release, described in Bundled examples and utilities.
Python tiler#
The tiler defines how the operator is partitioned across L3, L2, and L1 memory,
and across the 4x4 AIE core grid. For the full tiling reference, see
tiling.md in the bundled material described in
Bundled examples and utilities.
Vitis AI calls a getTiling() entry point:
def getTiling(opInterface: tensor_expr.OperatorInfo,
tiling: tensor_expr.AieConfig):
stamp = tiling[0][0] # Get the configurator for a stamp
# 1. Declare buffers at each memory level (TensorVar.make ...)
# 2. Tile tensors with .TileBy(index, size)
# 3. Configure data movement
# stamp.set_l3_to_l2_transfer(...)
# stamp.set_l2_to_l1_transfer(...)
# stamp.set_l1_to_l2_transfer(...)
# stamp.set_l2_to_l3_transfer(...)
# 4. Configure the kernel
# stamp.set_kernel_arguments([ifm_mk_var, wts_mk_var, ofm_mk_var])
# stamp.set_kernel_function_name("affinegrid2d_kernel")
# stamp.set_kernel_impl(Path("customop_kernel.cpp"))
# stamp.set_kernel_params([lp_params])
# stamp.set_kernel_nb_calls(num_ifm_tiles)
Registration#
Registration takes two independent steps, one for the compiler and one for the runtime. Both must agree on the operator’s domain and name.
Step |
Where |
Purpose |
|---|---|---|
Compiler |
|
Registers the operator under the |
Runtime |
Python, before session creation |
Declares the operator’s schema, that is its domain, name, and input and output counts, so that ONNX Runtime can parse the model graph. |
Compiler#
Register each operator under the vaiml_partition pass. The key takes the
form <domain>.<op_type_in_onnx>, and the value points to the YAML
configuration:
{
"passes": [
{
"name": "vaiml_partition",
"plugin": "vaip-pass_vaiml_partition",
"vaiml_config": {
"device": "ve2",
"custom_ops": {
"mydomain.myAffineGrid2D": {
"op_config": "customop_affinegrid2d/customop_affinegrid2d.yaml"
},
"mydomain.myScatterND": {
"op_config": "customop_scatternd/customop_scatternd.yaml"
},
"mydomain.myGridSample2D": {
"op_config": "customop_gridsample2d/customop_gridsample2d.yaml"
}
}
}
}
],
"target": "VAIML",
"targets": [
{
"name": "VAIML",
"pass": [
"init",
"vaiml_partition"
]
}
]
}
Runtime#
ONNX Runtime must know each operator’s schema before the inference session is
created. Register schemas with register_dynamic_custom_ops_to_onnxruntime().
This example keeps the schema information in a helper file,
custom_ops_config.json:
{
"mydomain.myAffineGrid2D": {
"domain": "mydomain",
"name": "myAffineGrid2D",
"inputs": 2,
"outputs": 1,
"op_config": "customop_affinegrid2d/customop_affinegrid2d.yaml"
}
}
Then register from it:
import json
from onnxruntime_custom_ops import (
register_dynamic_custom_ops_to_onnxruntime,
vaiml_custom_op_schema,
)
with open("custom_ops_config.json") as f:
custom_ops = json.load(f)
ops = [
vaiml_custom_op_schema(
domain=op["domain"],
name=op["name"],
nb_inputs=op["inputs"],
nb_outputs=op["outputs"],
)
for op in custom_ops.values()
]
register_dynamic_custom_ops_to_onnxruntime(ops)
Important
The registered name must match the operator type used inside the ONNX model exactly. A mismatch produces:
Fatal error: mydomain:<opname>(-1) is not a registered function/op.
Once the schemas are registered, create the inference session with
VitisAIExecutionProvider, setting config_file to
vitisai_config.json. The compiler partitions the graph, compiles each custom
operator subgraph according to its YAML configuration, and writes the compiled
artifacts to <cache_dir>/<cache_key>/.
Bundled examples and utilities#
Tutorials, worked examples, and supporting utilities ship in the release container, under:
/usr/local/lib/python3.12/dist-packages/flexml/flexml_extras/ai_utils
Location under |
Contents |
|---|---|
|
Complete worked custom operator examples, including |
|
Python utilities for compiling and simulating an operator, and for cutting and stitching ONNX graphs. Runnable directly from the command line. |
|
AIE architecture reference covering the constraints a custom operator must satisfy. |
|
Tiling reference covering memory partitioning and data movement. |