What is ROCm Compute Profiler?#

ROCm Compute Profiler is a kernel-level profiling tool for machine learning and high performance computing (HPC) workloads running on AMD Instinct™ GPUs. It measures how efficiently a single GPU kernel dispatch uses the underlying hardware: compute unit occupancy, cache and memory bandwidth utilization, instruction mix, and achieved versus peak arithmetic throughput.

This kernel-level view matters once you already know which kernel is worth optimizing. Where ROCm Systems Profiler characterizes an entire application to find where time goes, ROCm Compute Profiler drills into a specific dispatch to explain why it runs at the speed it does, down to the hardware block level.

AMD Instinct MI-Series GPUs are data center-class GPUs designed for compute and have some graphics capabilities disabled or removed. ROCm Compute Profiler primarily targets use with GPUs in the AMD Instinct MI300, MI200, and MI100 Series. Development is in progress to support Radeon™ (RDNA) GPUs.

ROCm Compute Profiler is built on top of ROCprofiler-SDK to monitor hardware performance counters. For the full performance-model reference, see Performance model.

How it works#

ROCm Compute Profiler is a Python-based tool that profiles an application in up to two stages: it first replays the application as many times as needed to collect the requested hardware counters per kernel dispatch, then, unless disabled with --no-roof, runs a set of accelerator-specific micro-benchmarks to establish the empirical roofline. The roofline model is not available on accelerators pre-MI200. See profiling-routine for the exact stage breakdown.

Diagram of ROCm Compute Profiler replaying an application across multiple counter-collection passes, running roofline micro-benchmarks, deriving metrics, and producing analysis views available as CLI output by default, or as CSV, an analysis database, or an HTML roofline plot

Counter collection is performed by injecting two libraries into the target process alongside ROCprofiler-SDK: a native tool (librocprofiler-compute-tool.so) that collects hardware performance counters per dispatch, and an SDK tool (librocprofiler-sdk-tool.so) that handles kernel tracing and output database generation. See Configuring the environment for profiling for the full backend and library breakdown.

Warning

Because associating counters with a specific kernel dispatch requires serializing dispatches, kernels launched on separate HIP streams on the same GPU don’t execute concurrently while profiling. Profiled kernel duration and utilization metrics reflect this serialized execution, not the concurrent behavior of a normal run. See the full warning and its consequences in Profile mode.

Once counters are collected, ROCm Compute Profiler derives higher-level metrics, such as utilizations and ratios, from the raw counter values to populate the analysis views described next.

ROCm Compute Profiler standalone GUI analyzer (experimental)#

ROCm Compute Profiler provides a standalone GUI to enable basic performance analysis.

Key features#

The feature set is documented across several how-to and conceptual pages; this table is a quick index into them.

Feature

Description

See also

Kernel-level counter collection

Collects hardware performance counters per kernel dispatch, replaying the application as many times as needed to gather every requested counter.

How it works

System and hardware block Speed-of-Light

Compares achieved utilization against theoretical peak, both for overall GPU throughput and for individual hardware blocks such as the command processor, workgroup manager, and caches.

Analysis views

Memory chart analysis

Visualizes transaction counts and bandwidth at each level of the memory hierarchy.

Analysis views

Empirical roofline analysis

Benchmarks achievable peak compute throughput and memory bandwidth, then plots kernel arithmetic intensity against that roofline.

Standalone roofline

Wavefront and instruction mix analysis

Reports launch and runtime statistics, and the breakdown of VALU, VMEM, scalar, and matrix instructions issued by a kernel.

Pipeline metrics

Baseline comparison

Compares two or more profiled runs of the same SoC side by side to check the effect of a code change.

analysis-baseline-comparison

Multiple analysis interfaces

The command line and ROCm Optiq both read the same profiling output.

Quickstart

Iteration multiplexing

Collects a large number of performance counters with minimal profiling overhead by splitting counter collection across kernel iterations instead of full application replays.

Iteration multiplexing

Analysis views#

ROCm Compute Profiler’s analysis report is organized into the following views:

View

What it tells you

See also

System Speed-of-Light

What percentage of the GPU’s theoretical peak performance the kernel achieves, across compute and memory subsystems.

CDNA Speed-of-Light, RDNA Speed-of-Light

Hardware block Speed-of-Light

Which specific hardware block — the compute unit, a cache level, the command processor, or the workgroup manager — is the bottleneck for a kernel, based on utilization and stall metrics for that block.

Analysis report block filtering, Performance model

Memory chart

Read, write, and atomic transaction counts, hit rates, and latencies at each level of the memory hierarchy: LDS, the L1 caches, the L2 cache, and HBM.

Vector L1 cache (vL1D), L2 cache (TCC)

Roofline

Classifies whether a kernel is compute-bound or memory-bound, and plots its exact position against attainable peak compute throughput and memory bandwidth; combine with kernel filtering for a per-kernel breakdown.

Standalone roofline, per-kernel-roofline

Wavefront analysis

Launch statistics (grid size, workgroup size, VGPR/AGPR/SGPR usage) and runtime statistics (wavefront occupancy, active cycles) for profiled kernels.

Pipeline metrics

Instruction mix

Breakdown of VALU, VMEM, scalar, LDS, and matrix (MFMA/WMMA) instructions issued by a kernel.

Pipeline metrics

Baseline comparison

Side-by-side comparison of any of the preceding views across two or more profiled runs of the same SoC — for example, before and after an optimization.

analysis-baseline-comparison

Key commands#

ROCm Compute Profiler exposes two primary modes on the rocprof-compute executable: profile to collect data, and analyze to read it back. See Modes for the full list of modes and Basic operations for the required arguments per operation.

The following are key profile and analyze commands:

  • To profile ./app, collecting all available counters for all kernels, and write results to workloads/my_run/<SoC>/:

    $ rocprof-compute profile -n my_run -- ./app
    
  • To profile only kernels whose name matches vecCopy:

    $ rocprof-compute profile -n my_run -k vecCopy -- ./app
    
  • To profile only the 1st, 2nd, and 3rd dispatch of each kernel:

    $ rocprof-compute profile -n my_run -d 1 2 3 -- ./app
    
  • To profile only the counters needed for specific analysis report blocks, which speeds up the profiling run:

    $ rocprof-compute profile -n my_run -b 2 5 -- ./app
    
  • To run only the roofline micro-benchmarks, skipping standard counter collection:

    $ rocprof-compute profile -n my_run --roof-only -- ./app
    
  • To analyze a profiled run in the terminal:

    $ rocprof-compute analyze -p workloads/my_run/MI300X/
    
  • To write the analysis to a SQLite analysis database instead of printing to the terminal:

    $ rocprof-compute analyze -p workloads/my_run/MI300X/ --output-format db
    
  • To compare two profiled runs of the same SoC side by side:

    $ rocprof-compute analyze -p run_a/MI300X/ run_b/MI300X/
    

For full walkthroughs, see Profile mode and CLI analysis.

Flags reference#

ROCm Compute Profiler supports many flags for filtering kernels, dispatches, and hardware blocks during profiling and analysis, and for selecting the analysis output format. See Profile mode for the full list of profile mode flags and Analyze mode for analyze mode flags, or run rocprof-compute profile -h or rocprof-compute analyze -h for the complete set.

Output formats#

analyze accepts an --output-format option for the analysis report. The following table lists the available values, plus the HTML roofline plots that analyze writes automatically, and the rocpd database that profile always writes:

Format

Produced by

Description

rocpd

profile

SQLite database of raw counters per dispatch, written automatically. Required for analyze --output-format csv or db.

stdout (default)

analyze

Analysis results printed to the terminal only.

txt

analyze

Analysis results written to a rocprof_compute_<uuid>.txt file.

csv

analyze

Analysis results written to a folder of CSV files. Requires the workload’s rocpd database from profiling.

db

analyze

Analysis results written to a rocprof_compute_<uuid>.db SQLite analysis database. Same requirement as csv.

html

analyze

Standalone Plotly roofline plot(s) (empirRoof_gpu-<id>...html), written automatically whenever the workload includes roofline data, independent of --output-format. See Standalone roofline.

Select the analyze output format with --output-format and override the file name with --output-name. See Analysis output format for details.

Supported hardware#

Roofline micro-benchmark support varies by SoC family independently of general counter and Speed-of-Light support. See Compatible GPUs/APUs for the full, tool-wide support table.

SoC family

Counters / SOL support

Roofline support

Notes

AMD Instinct™ MI350 Series (CDNA4, gfx950)

✅

✅

Full feature support; MI350P support introduced in ROCm Compute Profiler 3.6.0.

AMD Instinct MI300 Series (CDNA3)

✅

✅

Full feature support.

AMD Instinct MI200 Series (CDNA2)

✅

✅

Full feature support.

AMD Instinct MI100 Series (CDNA1)

✅

❌

Roofline micro-benchmarks require MI200 or later. See Standalone roofline.

AMD Ryzen™ AI Max / Ryzen AI Max+ 300 and 400 Series and Ryzen AI 5/7 PRO 4xx Series (Strix/Halo, Gorgon/Halo; gfx1150/gfx1151/gfx1152/gfx1153)

✅

✅

Full feature support; roofline benchmarking for gfx1150 and gfx1152 was added in ROCm Compute Profiler 3.7.0, with the gfx1151 (Strix Halo) gap closed in 3.8.0. Uses Wave Matrix Multiply Accumulate (WMMA) in place of MFMA.

AMD Instinct MI50, MI60 (Vega 20)

❌

❌

—