vaiprofile Utility#
vaiprofile is a tool available in the Docker container that converts profiling data from Vitis AI executions into per-layer statistics.
To use vaiprofile, first pre-process the profiling data generated from Vitis AI executions on the board, then run vaiprofile in the Docker environment.
Preprocess Trace Data#
After running inferences on the board, profiling data are generated and must be pre-processed and aggregated on the board before being transferred to the host for analysis with vaiprofile.
collect-profile-data <hardware_run_directory>
This script takes a hardware run directory, gathers the required top-level trace and timing JSON/TXT files, and packages them into a single ${run_dir}/vai_profile.zip archive for use with vaiprofile.
It collects all the files matching the following patterns: - aie_event_runtime_config_ctx_<hw_context_id>_run_<RunID>.json - aie_trace_ctx_<hw_context_id>_run_<RunID>_inf_<inferenceID>_strm_<StreamID>.txt - record_timer_ts*.json - record_timer_subgraph_cpu_ts.json - dtrace_dump_ctx_<timestamp>.py
It produces a vai_profile.zip in the specified <hardware_run_directory>.
Basic Profiling and DDR Bandwidth Analysis#
Use this tool to analyze trace dumps and DDR bandwidth data from Vitis AI executions. No additional flags are required: vaiprofile is data-driven and automatically extracts and processes the data available in the zip archive.
vaiprofile <VAIML_design_dir_path> vai_profile.zip [--frequency <freqMHz>]
vaiprofile --generate-profile-report <VAIML_design_dir_path> vai_profile_1.zip [vai_profile_2.zip ...] [--frequency <freqMHz>]
When used with a single profiling archive, vaiprofile generates per-layer statistics for that run and saves the results in the analyzed_data directory.
When used with --generate-profile-report and multiple archives, vaiprofile generates per-layer statistics for each individual run and a consolidated report combining results from all runs. Output is saved in the analyze_multiple_runs directory. For example, for 4 runs the directory structure is:
<VAIML_design_dir_path>/
└── analyze_multiple_runs/
├── 0/ # Processed profiling data for Run 0
├── 1/ # Processed profiling data for Run 1
├── 2/ # Processed profiling data for Run 2
├── 3/ # Processed profiling data for Run 3
└── merged_results/ # Detailed reports (use as logdir argument)
The directory analyze_multiple_runs/merged_results contains 4 .csv files and one .xlsx:
File |
Description |
|---|---|
|
Provides a summary of the model’s memory boundness for IFM and Weights, evaluated across both the L2–L1 and L3–L2 memory transfers. |
|
Provides a per-layer root-cause analysis of the identified model performance bottlenecks. |
|
Reports per-layer boundness cause, compute and stall cycles, and additional performance metrics averaged across the columns of the AI Engine Array. |
|
Extends |
|
Consolidates all of the preceding files into a single workbook, with each file presented on a separate tab. |