Runtime DDR Bandwidth Profiling Utility#
Purpose#
DDR bandwidth profiling configured via vitisai_config.json provides layer-by-layer bandwidth measurements at the AI Engine array/NoC interface. However, this approach does not identify which physical DDR device is the source of the observed traffic.
To complement this, Vitis AI provides two Tcl scripts that enable DDR bandwidth profiling at the DDRMC (DDR Memory Controller) level. This finer-grained analysis reveals the memory traffic breakdown per physical DDR device, allowing you to determine whether a bandwidth anomaly — previously identified through enhanced or detailed profiling — is attributable to a specific DDR device or reflects a broader memory traffic issue.
Scripts for DDR Bandwidth Profiling#
The scripts are available in Vitis AI: misc/versal_2ve/tests/ddr_bw_profiling/.
init_ddrmc_perf_counters.tcl
mc_read_ddr_bandwidth.tcl
init_ddrmc_perf_counters.tcl#
This script initializes the hardware performance counters within the DDR Memory Controller (DDRMC). It must be executed prior to any profiling session to ensure that the counters are correctly configured and ready to capture memory traffic data.
mc_read_ddr_bandwidth.tcl#
This script retrieves the DDR bandwidth measurements recorded by the DDRMC performance counters. It must be executed after init_ddrmc_perf_counters.tcl and while the target workload is running, in order to capture an accurate snapshot of memory traffic activity.
Methodology#
The DDR bandwidth profiling procedure consists of the following steps:
Connect to the target board via xsdb: Launch the Xilinx System Debugger (xsdb) on the host machine and establish a connection to the target board using the
connectcommand. This session serves as the execution environment for the subsequent profiling scripts.Initialize the DDRMC performance counters: Execute init_ddrmc_perf_counters.tcl within xsdb to configure the performance counters in the DDR Memory Controller. This step ensures that the hardware is ready to record memory traffic data before the workload begins.
Execute the model workload: Run the target model workload on the board to generate representative memory traffic. This step is required to produce the data that the DDRMC performance counters capture.
Collect DDR bandwidth measurements: While the workload is running, execute mc_read_ddr_bandwidth.tcl to sample the DDR bandwidth usage from the performance counters. The resulting data provides a per-channel breakdown of memory traffic, enabling precise identification of DDR devices contributing to bandwidth utilization.
Data Analysis#
Upon completion of the profiling session, an output.csv file is generated in the directory from which xsdb was launched. This file contains time-series DDR bandwidth measurements that can be used to characterize memory traffic patterns across the profiling interval.
The data may be analyzed using standard tools such as Microsoft Excel, Python (with pandas and matplotlib), or any CSV-compatible analysis software. This analysis supports the identification of memory bandwidth bottlenecks and informs optimization of memory access patterns.
Example output.csv#
DDR0_READ,DDR0_WRITE,DDR0_TOTAL,DDR1_READ,DDR1_WRITE,DDR1_TOTAL,DDR2_READ,DDR2_WRITE,DDR2_TOTAL,DDR3_READ,DDR3_WRITE,DDR
3_TOTAL,DDR4_READ,DDR4_WRITE,DDR4_TOTAL,TOTAL_READ,TOTAL_WRITE,TOTAL_BW
0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.01,0.0,0.01,0.0,0.0,0.0,0.01,0.0,0.01
0.01,0.01,0.02,0.01,0.01,0.02,0.01,0.0,0.01,0.01,0.01,0.02,0.76,1.17,1.93,0.8,1.2,2.0
0.06,0.03,0.09,0.02,0.02,0.04,0.01,0.11,0.13,0.0,0.0,0.01,0.18,0.02,0.2,0.27,0.18,0.47
0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0
0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.06,0.06,0.0,0.06,0.06
0.0,2.1,2.1,0.0,2.1,2.11,0.0,2.12,2.12,0.0,0.0,0.0,0.0,0.01,0.01,0.0,6.33,6.34
2.69,0.0,2.7,0.0,0.0,0.0,9.22,1.72,10.94,6.17,0.0,6.17,0.0,0.0,0.0,18.08,1.72,19.81
6.28,0.0,6.28,0.0,0.0,0.0,0.0,1.61,1.61,1.1,3.04,4.14,0.0,0.0,0.0,7.38,4.65,12.03
...
For each of the five DDR channels (DDR0 through DDR4), the output.csv file records per-sample read, write, and aggregate bandwidth values. This per-channel granularity enables targeted analysis of memory traffic distribution, facilitating the detection of bandwidth imbalances or saturation on individual DDR devices.