vaiprofile Utility#

vaiprofile is a tool available in the Docker container that converts profiling data from Vitis AI executions into per-layer statistics.

To use vaiprofile, first pre-process the profiling data generated from Vitis AI executions on the board, then run vaiprofile in the Docker environment.

Preprocess Trace Data#

After running inferences on the board, profiling data are generated and must be pre-processed and aggregated on the board before being transferred to the host for analysis with vaiprofile.

collect-profile-data <hardware_run_directory>

This script takes a hardware run directory, gathers the required top-level trace and timing JSON/TXT files, and packages them into a single ${run_dir}/vai_profile.zip archive for use with vaiprofile.

It collects all the files matching the following patterns: - aie_event_runtime_config_ctx_<hw_context_id>_run_<RunID>.json - aie_trace_ctx_<hw_context_id>_run_<RunID>_inf_<inferenceID>_strm_<StreamID>.txt - record_timer_ts*.json - record_timer_subgraph_cpu_ts.json - dtrace_dump_ctx_<timestamp>.py

It produces a vai_profile.zip in the specified <hardware_run_directory>.

Basic Profiling and DDR Bandwidth Analysis#

Use this tool to analyze trace dumps and DDR bandwidth data from Vitis AI executions. No additional flags are required: vaiprofile is data-driven and automatically extracts and processes the data available in the zip archive.

vaiprofile <VAIML_design_dir_path> vai_profile.zip [--frequency <freqMHz>]
vaiprofile --generate-profile-report <VAIML_design_dir_path> vai_profile_1.zip [vai_profile_2.zip ...] [--frequency <freqMHz>]

When used with a single profiling archive, vaiprofile generates per-layer statistics for that run and saves the results in the analyzed_data directory.

When used with --generate-profile-report and multiple archives, vaiprofile generates per-layer statistics for each individual run and a consolidated report combining results from all runs. Output is saved in the analyze_multiple_runs directory. For example, for 4 runs the directory structure is:

<VAIML_design_dir_path>/
└── analyze_multiple_runs/
    ├── 0/                    # Processed profiling data for Run 0
    ├── 1/                    # Processed profiling data for Run 1
    ├── 2/                    # Processed profiling data for Run 2
    ├── 3/                    # Processed profiling data for Run 3
    └── merged_results/       # Detailed reports (use as logdir argument)

The directory analyze_multiple_runs/merged_results contains 4 .csv files and one .xlsx:

Model Performance Output Files#

File

Description

model_performance_summary.csv

Provides a summary of the model’s memory boundness for IFM and Weights, evaluated across both the L2–L1 and L3–L2 memory transfers.

model_performance_cause.csv

Provides a per-layer root-cause analysis of the identified model performance bottlenecks.

model_performance_gist.csv

Reports per-layer boundness cause, compute and stall cycles, and additional performance metrics averaged across the columns of the AI Engine Array.

<model_name>.csv

Extends model_performance_gist.csv with micro-controller activity and detailed column-by-column metrics.

<model_name>.xlsx

Consolidates all of the preceding files into a single workbook, with each file presented on a separate tab.

Last updated on October 03, 2026.