ROCm Communication Collectives Library (RCCL) test configuration files#
2026-09-24
8 min read time
RCCL tests in CVS validate distributed GPU communication performance across AMD GPU clusters. The suites run RCCL collectives, optionally validate against expected thresholds, and generate HTML artifacts (graph and heatmap reports).
RCCL test suites#
CVS provides the following RCCL suites:
Test suite |
What it does |
|---|---|
|
User-facing performance suite that runs configured collectives with environment script staging. |
|
Regression suite with Cartesian product sweep using environment variable combinations from JSON configuration. |
All suites also collect host/network information and check firewall state before performance runs.
Run RCCL test commands#
See Run CVS RCCL performance and regression tests for cvs run examples, environment script staging, and heatmap generation.
Note
Update expected thresholds in results for your platform and cluster size before relying on pass/fail status.
Placeholders such as {user-id} in config paths are resolved by CVS at runtime.
rccl_config.json#
Main config file used by rccl_perf and rccl_regression:
rccl_config.json (example)
{
"rccl": {
"env_source_script": "/home/{user-id}/thor2_env_script.sh",
"mpi_params": {
"no_of_nodes": "2",
"no_of_local_ranks": "8",
"mpi_pml": "auto",
"mpi_dir": "/home/{user-id}/openmpi/bin",
"mpi_oob_port": "eth0"
},
"rccl_test_params": {
"rccl_collective": ["all_reduce_perf", "all_gather_perf", "broadcast_perf"],
"rccl_tests_dir": "/home/{user-id}/rccl-tests/build",
"start_msg_size": "1024",
"end_msg_size": "16g",
"step_function": "2",
"warmup_iterations": "10",
"no_of_iterations": "20",
"data_types": ["float"]
},
"cvs_params": {
"verify_bus_bw": "False",
"verify_bw_dip": "True",
"verify_lat_dip": "True",
"cluster_snapshot_debug": "False",
"rccl_result_file": "/home/{user-id}/rccl_result_file.json"
},
"results": {}
}
}
rccl_config.json with regression (example)
{
"rccl": {
"env_source_script": "/home/{user-id}/thor2_env_script.sh",
"mpi_params": {
"no_of_nodes": "2",
"no_of_local_ranks": "8",
"mpi_pml": "auto",
"mpi_dir": "/home/{user-id}/openmpi/bin"
},
"rccl_test_params": {
"rccl_collective": ["all_reduce_perf"],
"rccl_tests_dir": "/home/{user-id}/rccl-tests/build",
"start_msg_size": "1024",
"end_msg_size": "16g",
"step_function": "2",
"warmup_iterations": "10",
"no_of_iterations": "20"
},
"regression": {
"NCCL_ALGO": ["Ring", "Tree"],
"NCCL_PROTO": ["Simple"],
"NCCL_IB_QPS_PER_CONNECTION": ["1", "2"],
"NCCL_PXN_DISABLE": ["0", "1"],
"NCCL_MIN_NCHANNELS": ["8", "16"],
"NCCL_MAX_NCHANNELS": ["8", "16"]
},
"cvs_params": {
"verify_bus_bw": "False",
"verify_bw_dip": "True",
"verify_lat_dip": "True",
"rccl_result_file": "/home/{user-id}/rccl_result_file.json"
},
"results": {}
}
}
Parameters#
Configuration parameters for RCCL suites:
Core Configuration:
Configuration parameter |
Example/default |
Description |
|---|---|---|
|
|
Environment script sourced before test execution (contains RCCL/NCCL/UCX tuning parameters). |
|
|
RCCL collectives to execute. |
|
|
Output file path for RCCL parsed results. |
|
|
Start message size for sweep. |
|
|
End message size for sweep. |
|
|
Cartesian product sweep using NCCL/RCCL environment variable combinations (rccl_regression only). |
MPI Parameters (mpi_params):
Configuration parameter |
Example/default |
Description |
|---|---|---|
|
|
MPI point-to-point messaging layer (auto, ucx, ob1). |
RCCL Test Parameters (rccl_test_params):
Configuration parameter |
Example/default |
Description |
|---|---|---|
|
|
Warmup iteration count before measured iterations. |
|
|
Number of measured iterations. |
|
|
Message-size progression rule. |
CVS Parameters (cvs_params):
Configuration parameter |
Example/default |
Description |
|---|---|---|
|
|
Enable bus-bandwidth threshold validation. |
|
|
Enable bandwidth-dip validation. |
|
|
Enable latency-dip validation. |
|
|
Enables before/after cluster metric snapshots around tests. |
|
|
Expected threshold values used for pass/fail validation. |
Expected results format#
The results section is used for threshold validation. Values are keyed by collective and message size (bytes), with expected bus bandwidth values.
results snippet
"results": {
"all_reduce_perf": {
"bus_bw": {
"8589934592": "330.00",
"17179869184": "350.00"
}
}
}
Collective meanings#
The following collectives can be specified in the results block; each name maps to a distinct RCCL operation.
all_reduce_perf: all ranks reduce then receive the reduced result.all_gather_perf: each rank receives data from all ranks.scatter_perf: root rank distributes shards to all ranks.gather_perf: all ranks send data to a root rank.reduce_scatter_perf: reduction followed by scatter.sendrecv_perf: point-to-point pair communication.alltoall_perf: equal-sized all-to-all exchange.alltoallv_perf: variable-sized all-to-all exchange.broadcast_perf: one-to-all data broadcast.
Regression testing#
The rccl_regression suite performs Cartesian product sweeps using the regression object in the configuration. Each key-value pair represents an NCCL/RCCL environment variable and its possible values:
"regression": {
"NCCL_ALGO": ["Ring", "Tree"],
"NCCL_PROTO": ["Simple"],
"NCCL_IB_QPS_PER_CONNECTION": ["1", "2"],
"NCCL_PXN_DISABLE": ["0", "1"],
"NCCL_MIN_NCHANNELS": ["8", "16"],
"NCCL_MAX_NCHANNELS": ["8", "16"]
}
Key features:
Cartesian product: All combinations of the environment variables are tested
Paired channels:
NCCL_MIN_NCHANNELSandNCCL_MAX_NCHANNELSare paired (not Cartesian)Tree filtering: Tree algorithm only runs with
all_reduce_perfcollectiveEnvironment variables: Use actual NCCL/RCCL environment variable names as keys
Single-node testing#
Single-node RCCL testing can be achieved by configuring a single-node cluster in your cluster JSON file and running either rccl_perf or rccl_regression suites. The tests will automatically adapt to single-node execution when only one node is present in the cluster configuration.
Heatmap generation#
Generate performance heatmaps by comparing actual test results against a golden reference:
The following command generates an HTML heatmap that highlights per-collective bandwidth deviations from the reference.
cvs generate heatmap \
--actual /tmp/rccl_perf_results.json \
--reference /path/to/golden_reference.json \
--output /var/www/html/cvs/rccl_heatmap.html \
--title "RCCL Performance Comparison" \
--metadata
Options:
--actual: Path to actual test results JSON file (required)--reference: Path to golden reference JSON file (required)--output: Output HTML file path (optional, defaults to /tmp)--title: Custom heatmap title (optional)--metadata: Include metadata table if actual JSON has ‘metadata’ key--no-data-table: Exclude data table from output
Environment script setup#
All cluster configurations require an environment script to be sourced before RCCL tests, regardless of NIC type (Broadcom/ConnectX/AINIC):
Follow these steps to select and configure the correct environment script for your cluster hardware.
For AINIC clusters: Ensure AMD ANP is installed and available on all target nodes, then edit
input/config_file/rccl/ainic_env_script.shand setANP_HOME_DIRto your ANP install path.For other NIC types: Use the appropriate environment script (
thor2_env_script.shfor Broadcom,cx7_env_script.shfor ConnectX-7, etc.).Set
env_source_scriptin RCCL JSON config to the correct script path for your hardware.Run RCCL using
cvs run ...commands; CVS sources the script automatically.
The environment script contains essential RCCL/NCCL/UCX tuning parameters and paths required for proper test execution.
Known issue: bnxt_re (Thor2) multi-node GPU Direct RDMA failure#
Symptom: On bnxt_re (Broadcom Thor2) clusters, multi-node RCCL/NCCL jobs fail
symmetrically on every node during ncclCommInitRank with:
NET/IB: Peermem not available, falling back to GPU Direct RDMA (DMAbuf) for device 0
Call to ibv_reg_mr_iova2 failed with error Bad address
with a matching kernel-side error in dmesg:
infiniband bng_re0: bng_re_reg_user_mr: ib_umem_get failed! rc = -14
Root cause: This is not a per-node config or cabling issue – it reproduces
identically on every node because it is a build/config gap. RCCL’s
src/transport/net.cc only compiles its ROCm-native GPU memory registration
path (hsa_amd_portable_export_dmabuf) when built against
HIP_VERSION < 71260540. On newer ROCm builds (HIP_VERSION >= 71260540),
it instead takes a CUDA-style code path that additionally requires
ncclCuMemEnable() (NCCL_CUMEM_ENABLE) to be true. Because
NCCL_CUMEM_ENABLE defaults to 0 on ROCm/RCCL, that condition is never
met, so RCCL silently falls through to a plain (non-DMA-BUF) ibv_reg_mr_iova2
call on a raw GPU device pointer. Without a peer-memory kernel module,
bnxt_re’s driver cannot pin that pointer via get_user_pages and rejects
it with EFAULT (“Bad address”).
Fix: Set NCCL_CUMEM_ENABLE=1 (alongside NCCL_DMABUF_ENABLE=1) in the
environment script used for bnxt_re/Thor2 clusters. input/env_file/rccl/thor2_env_script.sh
sets this by default. Enabling CUMEM also lets NCCL recognize GPUs on a
UALoE/scale-up fabric as directly P2P-reachable (P2P/CUMEMMNNVL), bypassing
the NET/IB transport – and its DMA-BUF gap – entirely when a scale-up fabric
is present.
Verified on a 2-node bnxt_re/Helios cluster (4 GPUs/node, 8 ranks total): with
NCCL_CUMEM_ENABLE=0 every collective failed at ncclCommInitRank; with
NCCL_CUMEM_ENABLE=1 all collectives (all_reduce_perf, all_gather_perf,
reduce_scatter_perf, alltoall_perf, broadcast_perf; fp32 and bf16;
1 KB-1 GB) passed.
Validation and artifacts#
After a run, CVS produces the following outputs that can be used for pass/fail assessment and further analysis.
Test-level pass/fail is based on command execution plus enabled validations (
results, bandwidth checks, dip checks).Performance reports are generated under
/tmp/rccl_perf_report_*.html.Regression reports are generated under
/tmp/rccl_perf_report_*.htmlwith “RCCL Multi Node Performance Report” title.Heatmaps are generated using
cvs generate heatmapwith customizable output paths.All artifacts generated under
/tmpcan be copied to a web server path (for example/var/www/html/cvs) for browser access.