Using rocprofv3 on WSL2#
ROCprofiler-SDK has experimental support for running under WSL2 (Windows
Subsystem for Linux). On WSL2 the GPU is exposed through the DirectX paravirt
driver (/dev/dxg) and the DXCore user-mode library rather than the native
KFD driver (/dev/kfd). Because /dev/kfd is not present, ROCprofiler-SDK
enumerates agents differently and arms hardware counters through the DXG
vendor-packet path instead of the KFD profiling ioctl.
This page describes how to set up a working WSL2 profiling environment, what
rocprofv3 collects once it is set up, and which results should not be taken
at face value. The behavior described here was measured on a gfx1150
(AMD Radeon 890M) system running WSL2 with /dev/dxg and no /dev/kfd.
Other GPUs and other WSL2 installations are expected to behave the same way but
have not been measured.
Setting up#
Matching the runtime and the DXG thunk#
librocdxg.so (the DXG thunk) and libhsa-runtime64.so must be built from
the same source. This is a requirement, not a recommendation.
Important
HsaNodeProperties, the record the thunk fills in for each KMT node, is an
internal structure with no stable layout: it is 376 bytes in the ROCm 7.2.x
packages and 396 bytes in current development sources.
hsaKmtGetNodeProperties() carries no size argument, so the thunk writes
sizeof() bytes as it knows the type, into storage sized as the caller
knows the type. Pairing a newer thunk with an older runtime therefore
overruns the older side’s buffer and corrupts the heap, which typically
surfaces later as a crash somewhere unrelated.
The read is bounded by its destination rather than by the interface.
hsaKmtGetNodeProperties() copies the record whole and takes no size
parameter, so each node is read into storage deliberately longer than the
record this build expects, and a thunk that writes a longer one overruns that
slack instead of the rest of the stack frame. An optional DxgAbiCheck
handshake, by which a thunk reports the sizeof(HsaNodeProperties) it was
built with, is called when the thunk exports it, so whether the layout can be
negotiated depends on which librocdxg is installed: the released packages
export it, and a thunk reporting a different size is refused rather than read,
while a thunk built from the in-tree sources does not, and its records are read
on the assumption that the layouts match. A layout rearranged without a change
of sizeof() cannot be detected either way. None of this substitutes for a
matched pair.
Late attach has an additional runtime requirement. Released external
librocdxg packages do not reference-count the topology snapshot, so
ROCprofiler-SDK recognizes rocprofiler-register’s explicit late-attach marker
and disables WSL agent enumeration before loading or calling librocdxg.
Consequently, rocprof-attach is unsupported on WSL until a refcount-capable
runtime containing the #10034 snapshot ownership fix is installed and this
temporary restriction is removed. Ordinary rocprofv3 launch remains safe
with an older matched package because its topology read completes in a library
constructor before hsa_init(). This permits #7016 to land first and #10034
second. DxgAbiCheck is only the structure-size handshake described above;
its presence does not establish snapshot ownership.
Required environment variables#
Export the following before running rocprofv3:
# Resolve the matched runtime, thunk and ROCprofiler-SDK libraries ahead of
# any older system install under /opt/rocm
export LD_LIBRARY_PATH=<rocm-build>/lib:${LD_LIBRARY_PATH}
# Required with stock ROCm 7.2.x packages; already the default on current
# builds. Never set this to 0 on WSL2.
export HSA_ENABLE_DXG_DETECTION=1
Tool registration needs no variable of its own. A matched runtime offers its HSA
API table to rocprofiler-register on the DXG path exactly as it does on bare
metal, so nothing has to be exported to make a run profilable. A runtime that
still suppresses registration on the DXG path cannot be profiled at all: the
application runs to completion and rocprofv3 writes no output files — not
empty files, no files. That failure mode is silent, so if a run produces
nothing, confirm the runtime is the matched one described above.
HSA_TOOLS_DISABLE_REGISTER is the general HSA opt-out rather than something
WSL introduced, and it disables registration when it reads exactly 1. Leave
it unset; a 1 exported elsewhere in your environment for unrelated reasons
suppresses profiling output here just as it does anywhere else.
HSA_ENABLE_DXG_DETECTION is an HSA runtime variable rather than a
ROCprofiler-SDK one. On a current build the runtime enables DXG detection unless
the variable is explicitly set to 0, so exporting it is unnecessary but
harmless. With stock ROCm 7.2.x packages it is still opt-in, and without it
hsa_init() fails and every tool reports no supported GPU devices, which
reads like absent hardware rather than a configuration gap. Setting it to 0
disables the DXG path entirely and nothing on this page works.
Note
Export these from ~/.profile or from the job script rather than from
~/.bashrc. The default Ubuntu .bashrc returns early for
non-interactive shells, so scripts and CI would never see them.
Overriding librocdxg at run time#
A newer librocdxg can be placed ahead of the installed one on the loader
search path instead of replacing the system package, which is one way to get a
matched pair without touching /opt/rocm. Both the HSA runtime and
ROCprofiler-SDK ask for the unversioned name librocdxg.so, so a directory
holding that one symlink, placed first on LD_LIBRARY_PATH, redirects both
consumers:
mkdir -p ~/wsl-dxg-lib
ln -sf <path-to-matched-build>/librocdxg.so ~/wsl-dxg-lib/librocdxg.so
export LD_LIBRARY_PATH=~/wsl-dxg-lib:${LD_LIBRARY_PATH}
Warning
Preloading by absolute path — LD_PRELOAD=/path/to/librocdxg.so.<version>
— does not work, even though the preload itself succeeds. glibc registers
a preloaded object under the path it was given and under that object’s
SONAME. A request for librocdxg.so matches neither string, so the
loader falls through to a path search and quietly loads a second, stock copy
alongside the preloaded one. Preloading by bare name
(LD_PRELOAD=librocdxg.so) does work, because the loader resolves that
name through an ordinary search.
What rocprofv3 collects on WSL2#
The record counts in this section come from a HIP application that runs eight
iterations, each of which copies 1 MiB from host to device, launches a kernel,
and copies the result back. Run it under rocprofv3 the same way as on native
Linux:
rocprofv3 --sys-trace -d ./out -o results -- ./my_application
Option |
What was collected |
|---|---|
|
52 HIP API records, including exactly 16 |
|
Between 710 and 734 records across 38 distinct HSA functions, with no zero-length and no negative durations. |
|
8 dispatches, each with a kernel name, begin and end timestamps, VGPR and SGPR counts, and workgroup and grid dimensions. |
|
16 records: 8 host-to-device and 8 device-to-host. |
|
8 ranges named |
|
Four hardware counters collected together in one pass, with no PM4 command-buffer overflow. |
|
All of the above in one run, consistently with the individual options. |
Tracing and counter collection were also exercised together in four different combinations of options, and neither perturbed the other’s output.
Thread trace (ATT)#
--att works on WSL2. On native Linux, GPUs that support it take the thread
trace buffers, signals and copies directly from KFD; without /dev/kfd
ROCprofiler-SDK takes them from the HSA runtime instead, which needs no
configuration. Restricting the capture to ROCTx-selected regions with
--selected-regions produced empty traces on WSL2 in testing.
Decoding the captured data needs the rocprof-trace-decoder library, which is
distributed separately. Point rocprofv3 at the directory that holds
librocprof-trace-decoder.so with --att-library-path or the
ROCPROF_ATT_LIBRARY_PATH environment variable. When building the
ROCprofiler-SDK tests and samples against a decoder that is not installed under
${ROCM_PATH}, set ROCPROFILER_SDK_TRACE_DECODER_ROOT to it; the thread
trace tests and samples are skipped when no decoder is found.
Timestamps#
Timestamps are usable for building a timeline. Rather than inspecting the values
directly, correctness was checked by containment: every GPU dispatch falls
entirely inside the host interval between its hipLaunchKernel and the
following device-to-host hipMemcpy that consumes its result, and every copy
falls entirely inside the hipMemcpy that issued it. A skewed or wrongly
scaled GPU clock would break that relationship.
The reference workload has no explicit hipDeviceSynchronize; the blocking
copy is what synchronizes it. Where a workload does call one, the dispatch is
contained by the launch-to-synchronize interval instead, as a variant of the
same workload confirmed.
The host timeline is the WSL VM’s CLOCK_MONOTONIC in nanoseconds.
HSA_SYSTEM_INFO_TIMESTAMP_FREQUENCY reports 1 GHz, the frequency measures
999.9978 MHz, and the timeline agrees with CLOCK_MONOTONIC to within 175 ns.
The GPU agent’s own wall clock runs at 100 MHz (WallClockKHz = 100000) and
the runtime converts it, so agent-side and host-side timestamps can be compared
directly.
Hardware counter accuracy#
Counter values on WSL2 need to be interpreted with care. The SQ counters are global to the shader engines with no per-process scoping, and WDDM time-slices the GPU among all its clients, so any GPU work performed for another client inside a dispatch’s counting window is attributed to that dispatch. The Windows desktop compositor is one such client.
Apart from the zero read-backs described under Known issues below, the error is one-directional: measured values are never below the analytically expected result, and the smallest value observed over repeated runs is the one closest to the truth:
A 262,144 work-item kernel (wave32, so 8,192 waves in theory) reported a minimum
SQ_WAVESof exactly 8,192 in each of three different collection configurations. Maxima in the same runs reached 185,561.A 1,048,576 work-item kernel (32,768 waves in theory) reported 36,537, 111,929, 162,155, 260,666, 315,977 and 394,836 across six runs. The smallest of those is 11 percent above theory; none is below it.
The collector itself is therefore behaving correctly — where the floor is reached it lands exactly on the analytically derived value — and the excess is contamination from concurrent GPU work, not a counting error. What varies is how much foreign activity happened to land in the window.
To get usable numbers, sample the same workload several times and take the smallest non-zero value, or collect with the Windows desktop idle so there is little foreign work to absorb. Short-running kernels are affected disproportionately, because a fixed amount of contamination is a much larger fraction of a small count. ROCprofiler-SDK cannot correct for this: attributing counts back to a single process would require the hardware counters to be filtered per VMID, below the SDK.
Known issues#
GRBM counters read zero at the start of a session#
GRBM_* counters can read back zero on the opening dispatches of a profiling
session, on dispatches that demonstrably ran and produced waves. Such a zero is
invalid data rather than a measurement, and it matters because derived metrics
divide by these counters: GPU_UTIL and CU_UTILIZATION divide by
GRBM_COUNT, and VALUBusy, SALUBusy, MemUnitStalled,
LDSBankConflict and MeanOccupancyPerCU depend on GRBM_GUI_ACTIVE.
Discard the affected dispatches rather than letting the zero propagate into a
metric.
When GRBM-derived metrics matter, discard the first 300 ms of dispatches in the session, which covers the longest burst observed, or give the application a short warm-up phase so that the affected window falls before the region of interest.
The behavior is narrow and well characterized. It is confined to the start of a
session: across 354 sessions and 23,990 dispatches, 11 sessions (3.1 percent)
and 191 dispatches (0.80 percent) were affected, all 11 bursts began at the very
first dispatch of their session, and none began mid-session, including through
36 seconds of continuous profiling. Each burst is contiguous and does not
recur — once a GRBM counter reports a non-zero value it keeps reporting for the
rest of that session — and bursts ran from 9 ms to 291 ms, with a median around
54 ms. Every GRBM_* counter in a set fails together, so one GRBM counter
reading zero while another reads normally is not a case you have to handle.
The fault sits in the GRBM block itself: the counting window, the perfmon arming, the packet execution and the read-back have all been verified working, and during a burst the block is correctly armed and correctly programmed for the counter requested, indistinguishable from a healthy dispatch. What is left is a counter that does not advance, and why it does not is still unknown. It is at least not a matter of some other consumer of GPU utilization telemetry claiming the block — profiling alongside a Windows utilization monitor that samples continuously produced no affected sessions at all, the opposite of what competition for the counter slots would predict.
Nothing under your control changes the odds: the rate does not vary with the number or combination of counters requested, with dispatch duration, with a dispatch’s position in the session, or with how long the GPU sat idle beforehand. It does drift over time, though, and affected sessions arrive in clusters rather than independently — across roughly an hour of continuous profiling it moved between zero and about one session in ten. A short run that sees no zeros therefore does not establish that a machine is unaffected, and one that sees several does not mean it is unusually bad.
SQ counters are not completely immune to a zero read-back — 0.10 percent of dispatches showed one — but those are always isolated single dispatches rather than an opening burst, so the GRBM behavior is a distinct effect.
Behavior that is not a defect#
Two observations are easy to mistake for WSL2 gaps and are worth stating explicitly.
The memory copy CSV has no byte-count column. Its columns are Kind,
Direction, Stream_Id, Source_Agent_Id, Destination_Agent_Id,
Correlation_Id, Start_Timestamp and End_Timestamp. That is the
upstream schema on every platform, not something WSL2 drops. Byte counts are
available from the rocpd output, in rocpd_memory_copy.size:
rocprofv3 --memory-copy-trace --output-format rocpd -d ./out -o results -- ./my_application
Device-to-device copies within the same device do not appear in the memory copy
trace. HIP implements them as a blit kernel rather than as an SDMA copy, so they
appear in the kernel trace as __amd_rocclr_copyBuffer. Copies that cross
agents are unaffected.
Not covered on WSL2#
PC sampling and SPM are not covered by this page and were not verified in this configuration. Those paths reach the GPU through KFD interfaces that WSL2 does not expose, and ROCprofiler-SDK gates them on the presence of the KFD device.
How agents are discovered#
Agent discovery on WSL2 combines two sources, both read before any agent record
becomes visible through rocprofiler_query_available_agents():
DXCore (
libdxcore.so) enumerates the GPU adapters and supplies the adapter identity, its marketing name and the dedicated VRAM size.librocdxg (the DXG thunk) supplies the KMT node topology — compute unit and SIMD counts, shader engines and arrays, wavefront size, clocks, family, firmware versions and the gfx target — through the same interface the HSA runtime itself uses on
/dev/dxgsystems.
ROCprofiler-SDK loads librocdxg.so at run time with dlopen, reusing the
copy the HSA runtime already loaded when there is one, so no build-time
configuration and no link dependency are involved. Records are built once, at
enumeration time, and are never modified afterwards, so pre-HSA consumers such
as rocprofv3-avail see the same complete records as a profiling run.
Each adapter is paired with its KMT node by Windows LUID. The PCI device id is only a fallback for nodes that report no LUID, and it has to identify exactly one of them: two identical GPUs that report no LUID cannot be told apart, so both are omitted with a diagnostic rather than one being guessed at. Each node is claimed by at most one adapter. An adapter with no matching node, or whose node reports an incomplete topology, is likewise logged and omitted rather than published with placeholder values.
The read uses the thunk’s ordinary KMT entry points — hsaKmtOpenKFD,
hsaKmtAcquireSystemProperties, hsaKmtGetNodeProperties,
hsaKmtReleaseSystemProperties and hsaKmtCloseKFD. All of them must
resolve before any of them is called, so a thunk missing one is refused, with
the missing symbol named, before anything is opened.
When agents are missing#
A librocdxg that does not export what the read needs, or that offers the
DxgAbiCheck handshake and reports a different record layout, yields no GPU
agents at all in ROCprofiler-SDK. The log separates the two: a missing entry
point is named as such, while a rejected layout reads wsl topology: ...
rejected the HsaNodeProperties layout. Only a thunk offering the handshake can
be turned away for its layout; one that does not offer it is read on the
assumption that the layouts match. This is visible rather than silent, but it is
not harmless: the HSA runtime loads the
same thunk and still reports its GPUs, so ROCprofiler-SDK and HSA disagree about
which agents exist. Partially described topologies land in the same place — the
GPU that could be described is published and the other is omitted, while HSA
keeps reporting both.
Agent records are paired with HSA agents by the KMT node id the thunk reported, never by a dense ordinal, so the GPUs that were published stay correctly paired no matter which ones were omitted. HSA GPUs left without a counterpart are reported as an unsupported-profiling condition naming their KMT node ids: profiling is unavailable on those GPUs, the ones that did pair remain fully profilable, and the application under test keeps running. The log names the symbol or adapter that was rejected; the fix is to bring the WSL ROCm runtime package and ROCprofiler-SDK back into a matched pair.
Environment variables#
HSA_TOOLS_DISABLE_REGISTERAn existing HSA runtime variable that disables tool registration when it reads
1, leaving rocprofiler-register without an HSA API table to intercept. It takes no WSL2-specific value and should be left unset; set to1, here as anywhere else,rocprofv3produces no output files at all.HSA_ENABLE_DXG_DETECTIONAn HSA runtime variable. Current builds enable DXG detection unless this is explicitly set to
0; stock ROCm 7.2.x packages require1to be set, and otherwise failhsa_init()and reportno supported GPU devices. Setting it to0disables the DXG path on any build.
WSLKMT_VENDOR_PACKETGates whether the libhsakmt DXG path honors the vendor-specific PM4 IB packets that AQLprofile emits for hardware counter collection. When unset (the libhsakmt default), the embedded PM4 IB is silently dropped and every
Counter_Valuereads back zero. ROCprofiler-SDK sets it to1automatically when it detects a WSL GPU environment, using a non-overwriting write so an explicit user setting always wins. You normally do not need to set it; exportWSLKMT_VENDOR_PACKET=0to opt out.ROCPROFILER_FORCE_GFXOverrides the
gfxtarget name used to look up counter definitions (config.yamlis keyed by gfx target). The target normally comes from the KMT node’s engine id, which already honors the HSA runtime’s ownHSA_OVERRIDE_GFX_VERSION; set this variable to override both. The value must be of the formgfx<NNN>with at least three digits, the last of which may be a lowercase hex digit (as ingfx90a); a malformed value is ignored with a warning and the node-reported target is used.ROCPROFILER_FORCE_PLATFORMOverrides platform autodetection, which otherwise selects the KFD enumerator when KFD sysfs is present and the WSL enumerator when
/dev/dxgandlibdxcore.soare. Acceptsgnulinuxorwslon Linux; a value not built into the binary is logged and ignored. This is a diagnostic escape hatch and is not needed in normal use.