Using PC sampling in ROCm Compute Profiler#
Warning
PC sampling is an experimental feature. Enable it in profile
mode by passing --experimental --pc-sampling. The analyze
command detects PC sampling automatically from the profiling
configuration and needs no extra pc-sampling flag. Behavior and
command-line surface may change in future releases.
Program Counter (PC) sampling service for GPU profiling is a profiling technique that periodically samples the program counter during the GPU kernel execution to understand code execution patterns and hotspots.
ROCm Compute Profiler supports Host Trap PC sampling and Stochastic (Hardware-Based) PC sampling. Host Trap PC sampling is enabled for AMD Instinct MI200 Series and later GPUs. Stochastic (hardware-based) PC sampling is enabled for AMD Instinct MI300 Series and later GPUs. Stochastic PC sampling provides additional information that tells whether a sampled wave issued an instruction for a particular PC. It also provides the reason for not issuing the instruction (stall reason). This type of information is particularly useful for understanding stalls during the kernel execution. The PC sampling can be used with profiling and analysis options.
Profiling options#
For using profiling options for PC sampling the configuration needed are:
--pc-sampling-method: Should be eitherstochasticorhost_trap, (DEFAULT: stochastic)--pc-sampling-interval: The accepted range is read from the device; seerocprofv3-avail info --pc-sampling. When the device cannot be queried, 1 to 1048576 is accepted. Forstochasticsampling, the interval is in cycles and must be a power of 2 (DEFAULT: 1048576). Forhost_trapsampling, the interval is in microseconds (DEFAULT: 512). When omitted, the method-appropriate default is used.
Sample command:
$ rocprof-compute profile -n pc_test --no-roof --experimental --pc-sampling --pc-sampling-method stochastic -VVV -- target_app
Profile multi-process workloads#
The same profile command supports applications that create multiple processes. Every process that runs GPU kernels is sampled, and no additional option is required.
Stall reasons#
The stall_reason column of a stochastic PC sampling table holds the raw
reason names reported by ROCprofiler-SDK for samples where a wave did not
issue an instruction. Those names are defined in the ROCprofiler-SDK
documentation, and the analysis output prints a link to it above the table:
--------------------------------------------------------------------------------
21. PC Sampling
21.1 PC Sampling
Stall reason definitions: https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/how-to/cdna3-cdna4-pc-sampling.html#stall-reasons
Analysis options#
For using analysis options for PC sampling the configuration needed are:
--pc-sampling-sorting-type:offsetorcount. The default option iscount, which surfaces the most-sampled instructions (hotspots) first.offsetis an assembly instruction offset in the code object.--pc-sampling-rows: Maximum number of rows shown in the PC sampling table (DEFAULT: 10). Must be a non-negative integer; use0to show all rows.
Two columns describe how well each instruction used the machine:
active_thread_percent: percent of the wave’s lanes that were active. A low value means control flow masked lanes off. Shown for both sampling methods.wave_occupancy_percent: percent of the machine’s wave slots that held a wave. A low value means there were too few waves to hide latency. Shown forstochasticonly, becausehost_traprecords carry no wave count.
Both are averaged over the samples that landed on the instruction. A kernel
built for wave64 on a device that reports a wave size of 32 sets more lanes
than that size, so active_thread_percent is capped at 100 and analyze warns
once with how many lines were capped.
Sample command:
$ rocprof-compute analyze -p <workload_dir> -k 0 --pc-sampling-sorting-type offset
Sample output:
source_line shows N/A because the example binary was built without
-g (See the note at the end of this page).
instruction_type names the hardware pipeline that runs the instruction,
such as VALU for vector math or MATRIX for matrix math.
Selecting a single kernel with host_trap PC sampling:
$ rocprof-compute analyze -p <workload_dir> -k 0
╒═════════╤═════════╤═══════════════╤════════════════════════════════╤════════════════════╤══════════════════╤══════════╤═════════╤═════════════════════════╤══════════════════════════════════════╕
│ index │ pid │ source_line │ instruction │ instruction_type │ code_object_id │ offset │ count │ active_thread_percent │ Kernel_Name │
╞═════════╪═════════╪═══════════════╪════════════════════════════════╪════════════════════╪══════════════════╪══════════╪═════════╪═════════════════════════╪══════════════════════════════════════╡
│ 83 │ 1429079 │ N/A │ v_add_u32_e32 v16, s2, v0 │ VALU │ 2 │ 0x3f30 │ 42959 │ 100.00 │ matmul_fp32_throughput(float*, float │
│ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
├─────────┼─────────┼───────────────┼────────────────────────────────┼────────────────────┼──────────────────┼──────────┼─────────┼─────────────────────────┼──────────────────────────────────────┤
│ 84 │ 1429079 │ N/A │ v_ashrrev_i32_e32 v17, 31, v16 │ VALU │ 2 │ 0x3f38 │ 10908 │ 96.88 │ matmul_fp32_throughput(float*, float │
│ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
├─────────┼─────────┼───────────────┼────────────────────────────────┼────────────────────┼──────────────────┼──────────┼─────────┼─────────────────────────┼──────────────────────────────────────┤
│ 85 │ 1429079 │ N/A │ s_load_dword s0, s[0:1], 0x0 │ SCALAR │ 2 │ 0x3f40 │ 1 │ 87.50 │ matmul_fp32_throughput(float*, float │
│ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
╘═════════╧═════════╧═══════════════╧════════════════════════════════╧════════════════════╧══════════════════╧══════════╧═════════╧═════════════════════════╧══════════════════════════════════════╛
Selecting a single kernel with stochastic PC sampling, which adds the
count_issued, count_stalled, wave_occupancy_percent, and
stall_reason columns:
$ rocprof-compute analyze -p <workload_dir> -k 0
╒═════════╤═════════╤═══════════════╤════════════════════════════════════╤════════════════════╤══════════════════╤══════════╤═════════╤═════════════════════════╤════════════════╤═════════════════╤══════════════════════════╤═══════════════════════════════════════════════════════╤══════════════════════════════════════╕
│ index │ pid │ source_line │ instruction │ instruction_type │ code_object_id │ offset │ count │ active_thread_percent │ count_issued │ count_stalled │ wave_occupancy_percent │ stall_reason │ Kernel_Name │
╞═════════╪═════════╪═══════════════╪════════════════════════════════════╪════════════════════╪══════════════════╪══════════╪═════════╪═════════════════════════╪════════════════╪═════════════════╪══════════════════════════╪═══════════════════════════════════════════════════════╪══════════════════════════════════════╡
│ 90 │ 1429079 │ N/A │ s_load_dword s8, s[0:1], 0x24 │ SCALAR │ 2 │ 0x3f00 │ 3 │ 100.00 │ 0 │ 3 │ 78.12 │ [('ARBITER_NOT_WIN', 3)] │ matmul_fp32_throughput(float*, float │
│ │ │ │ │ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
├─────────┼─────────┼───────────────┼────────────────────────────────────┼────────────────────┼──────────────────┼──────────┼─────────┼─────────────────────────┼────────────────┼─────────────────┼──────────────────────────┼───────────────────────────────────────────────────────┼──────────────────────────────────────┤
│ 91 │ 1429079 │ N/A │ s_load_dwordx4 s[4:7], s[0:1], 0x0 │ SCALAR │ 2 │ 0x3f08 │ 2 │ 100.00 │ 0 │ 2 │ 71.88 │ [('ARBITER_NOT_WIN', 1), ('ARBITER_WIN_EX_STALL', 1)] │ matmul_fp32_throughput(float*, float │
│ │ │ │ │ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
├─────────┼─────────┼───────────────┼────────────────────────────────────┼────────────────────┼──────────────────┼──────────┼─────────┼─────────────────────────┼────────────────┼─────────────────┼──────────────────────────┼───────────────────────────────────────────────────────┼──────────────────────────────────────┤
│ 92 │ 1429079 │ N/A │ s_load_dword s3, s[0:1], 0x10 │ SCALAR │ 2 │ 0x3f10 │ 2 │ 96.88 │ 0 │ 2 │ 71.88 │ [('ARBITER_WIN_EX_STALL', 1), ('ARBITER_NOT_WIN', 1)] │ matmul_fp32_throughput(float*, float │
│ │ │ │ │ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
╘═════════╧═════════╧═══════════════╧════════════════════════════════════╧════════════════════╧══════════════════╧══════════╧═════════╧═════════════════════════╧════════════════╧═════════════════╧══════════════════════════╧═══════════════════════════════════════════════════════╧══════════════════════════════════════╛
Without a kernel filter, the same per-instruction table is shown across all
kernels, with a Kernel_Name column identifying each row’s kernel:
$ rocprof-compute analyze -p <workload_dir>
╒═════════╤═════════╤═══════════════╤══════════════════════════════════════════════════════════════╤════════════════════╤══════════════════╤══════════╤═════════╤═════════════════════════╤════════════════╤═════════════════╤══════════════════════════╤═══════════════════════════════════════════════╤══════════════════════════════════════════╕
│ index │ pid │ source_line │ instruction │ instruction_type │ code_object_id │ offset │ count │ active_thread_percent │ count_issued │ count_stalled │ wave_occupancy_percent │ stall_reason │ Kernel_Name │
╞═════════╪═════════╪═══════════════╪══════════════════════════════════════════════════════════════╪════════════════════╪══════════════════╪══════════╪═════════╪═════════════════════════╪════════════════╪═════════════════╪══════════════════════════╪═══════════════════════════════════════════════╪══════════════════════════════════════════╡
│ 12 │ 1429079 │ N/A │ s_cmp_eq_u32 s1, 0 │ SCALAR │ 2 │ 0x3bb8 │ 199 │ 100.00 │ 197 │ 2 │ 84.38 │ [('OTHER_WAIT', 197), ('ARBITER_NOT_WIN', 2)] │ _Z22matmul_fp16_throughputPDv4_DF16_PDv4 │
│ │ │ │ │ │ │ │ │ │ │ │ │ │ _fi │
├─────────┼─────────┼───────────────┼──────────────────────────────────────────────────────────────┼────────────────────┼──────────────────┼──────────┼─────────┼─────────────────────────┼────────────────┼─────────────────┼──────────────────────────┼───────────────────────────────────────────────┼──────────────────────────────────────────┤
│ 13 │ 1429079 │ N/A │ s_waitcnt vmcnt(2) │ INTERNAL │ 2 │ 0x3bbc │ 199 │ 100.00 │ 0 │ 199 │ 84.38 │ [('WAITCNT', 199)] │ _Z22matmul_fp16_throughputPDv4_DF16_PDv4 │
│ │ │ │ │ │ │ │ │ │ │ │ │ │ _fi │
├─────────┼─────────┼───────────────┼──────────────────────────────────────────────────────────────┼────────────────────┼──────────────────┼──────────┼─────────┼─────────────────────────┼────────────────┼─────────────────┼──────────────────────────┼───────────────────────────────────────────────┼──────────────────────────────────────────┤
│ 14 │ 1429079 │ N/A │ v_mfma_f32_16x16x16_f16 v[8:11], v[20:21], v[20:21], v[8:11] │ MATRIX │ 2 │ 0x3bc0 │ 1 │ 93.75 │ 0 │ 1 │ 62.50 │ [('ARBITER_NOT_WIN', 1)] │ void fma_throughput<int>(int │
│ │ │ │ │ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
╘═════════╧═════════╧═══════════════╧══════════════════════════════════════════════════════════════╧════════════════════╧══════════════════╧══════════╧═════════╧═════════════════════════╧════════════════╧═════════════════╧══════════════════════════╧═══════════════════════════════════════════════╧══════════════════════════════════════════╛
Sorting a single kernel by sample count instead of offset:
$ rocprof-compute analyze -p <workload_dir> -k 0 --pc-sampling-sorting-type count
╒═════════╤═════════╤═══════════════╤═══════════════════════════════════════════════════╤════════════════════╤══════════════════╤══════════╤═════════╤═════════════════════════╤════════════════╤═════════════════╤══════════════════════════╤══════════════════════════════════════════════════════════════╤══════════════════════════════════════╕
│ index │ pid │ source_line │ instruction │ instruction_type │ code_object_id │ offset │ count │ active_thread_percent │ count_issued │ count_stalled │ wave_occupancy_percent │ stall_reason │ Kernel_Name │
╞═════════╪═════════╪═══════════════╪═══════════════════════════════════════════════════╪════════════════════╪══════════════════╪══════════╪═════════╪═════════════════════════╪════════════════╪═════════════════╪══════════════════════════╪══════════════════════════════════════════════════════════════╪══════════════════════════════════════╡
│ 106 │ 1429079 │ N/A │ global_load_dword v18, v[0:1], off │ FLAT │ 2 │ 0x3f78 │ 29715 │ 100.00 │ 0 │ 29715 │ 90.62 │ [('ARBITER_NOT_WIN', 26037), ('ARBITER_WIN_EX_STALL', 3678)] │ matmul_fp32_throughput(float*, float │
│ │ │ │ │ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
├─────────┼─────────┼───────────────┼───────────────────────────────────────────────────┼────────────────────┼──────────────────┼──────────┼─────────┼─────────────────────────┼────────────────┼─────────────────┼──────────────────────────┼──────────────────────────────────────────────────────────────┼──────────────────────────────────────┤
│ 117 │ 1429079 │ N/A │ v_mfma_f32_16x16x4_f32 v[8:11], v19, v19, v[8:11] │ VALU │ 2 │ 0x3fb0 │ 21164 │ 100.00 │ 169 │ 20995 │ 87.50 │ [('ARBITER_NOT_WIN', 20995), ('OTHER_WAIT', 169)] │ matmul_fp32_throughput(float*, float │
│ │ │ │ │ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
├─────────┼─────────┼───────────────┼───────────────────────────────────────────────────┼────────────────────┼──────────────────┼──────────┼─────────┼─────────────────────────┼────────────────┼─────────────────┼──────────────────────────┼──────────────────────────────────────────────────────────────┼──────────────────────────────────────┤
│ 188 │ 1429079 │ N/A │ global_store_dwordx4 v[4:5], v[0:3], off │ FLAT │ 2 │ 0x4204 │ 13821 │ 78.12 │ 0 │ 13821 │ 65.62 │ [('ARBITER_WIN_EX_STALL', 7300), ('ARBITER_NOT_WIN', 6521)] │ matmul_fp32_throughput(float*, float │
│ │ │ │ │ │ │ │ │ │ │ │ │ │ __vector(4)*, int) │
╘═════════╧═════════╧═══════════════╧═══════════════════════════════════════════════════╧════════════════════╧══════════════════╧══════════╧═════════╧═════════════════════════╧════════════════╧═════════════════╧══════════════════════════╧══════════════════════════════════════════════════════════════╧══════════════════════════════════════╛
Analyze multi-process workloads#
Pass the workload directory to analyze as usual. No additional
multi-process option is required:
$ rocprof-compute analyze -p <workload_dir>
Analyze mode loads the PC sampling data for every process in the workload
directory and reports it in a single table. A pid column identifies the
process each row came from.
Samples are not merged across processes. The same instruction offset can belong to different code in two processes, so each process keeps its own rows and its own counts.
Sorting and --pc-sampling-rows apply to the combined table. Sorting by
offset orders on pid first, which keeps each process’s rows together.
Sorting by count ranks on sample count alone, so rows from different
processes can interleave. In both cases, --pc-sampling-rows 10 selects ten
rows from the workload as a whole, not ten rows per process.
Per-kernel ISA and source in CSV output#
--output-format csv writes, alongside the analysis tables, one file per
kernel per code object per process holding that kernel’s instruction lines with
the samples collected on them, and exports the source those instructions were
compiled from beside it:
<output_name>/
kernel.csv, pc_sampling_summary.csv, ...
per_kernel_pc_sampling/
<workload_name>/<workload_sub_name>/
source/<source path with the leading separator dropped>
<short_name>_uuid_<kernel_uuid>/
isa_code_object_id_<code_object_id>_pid_<pid>.csv
A folder leads with the kernel’s short name, the identifier its C++ signature
demangles down to, so vecCopy_2(double*, double*, double*, int, int) is
filed under vecCopy_2_uuid_7. Overloads and template instantiations share a
short name, so the kernel_uuid suffix keeps each folder unique. kernel.csv
carries kernel_uuid alongside short_name and kernel_name, so it maps
a folder back to the kernel it holds. One kernel has more than one file when it
was compiled into several code objects, or when several processes ran it.
Each file holds one row per instruction line, ordered by code object offset:
Column |
Description |
|---|---|
|
Position of the line within this file, counting from 1. |
|
Offset of the instruction from the code object’s load address. |
|
Disassembled instruction. |
|
Hardware pipeline that runs the instruction, such as |
|
Samples that landed on this instruction. |
|
Samples where the wave issued. Empty for |
|
Samples where the wave was stalled. Empty for |
|
Percent of the machine’s wave slots that held a wave at this
instruction. Empty for |
|
Percent of the wave’s lanes that were active at this instruction. A low value means control flow masked lanes off. |
|
Samples stalled at this instruction for that reason. |
|
Inline stack of |
|
Code object the instruction belongs to, local to its process. |
|
Process the code object was loaded in. |
The Stall <REASON> columns are the reasons the workload’s samples actually
carry, so they vary between runs, and a host_trap workload has none of them.
Every file of one workload has the same columns. A cell is empty where that
reason was not seen at that offset.
An instruction that no sample landed on keeps its row, with empty counts. It
still carries an Instruction type.
The Source column and the source/ folder are populated only when the
target app was built with debug info; see the note.
Note
PC sampling now only shows assembly instructions collected in our record of pc samples and not all instructions of compiled code are represented.
Source information requires the target app to be built with debug info (for example
hipcc -g). Without it, samples map to assembly only: profile mode captures no source files, the terminal table’ssource_lineshowsN/A, and theSourcecolumn of the per-kernel CSV is empty.The
Sourcecolumn can show?instead of a line number, such askernel.cpp:?. The compiler does this when it builds one instruction from several source lines. To find the source of such an instruction, look at the source lines of the instructions it depends on. For example, a wait instruction such ass_waitcntdepends on earlier memory and LDS operations, such asglobal_*,buffer_*,s_load*ords_*instructions.