ROCm Communication Collectives Library (RCCL) test configuration files#

2026-09-24

8 min read time

Applies to Linux

RCCL tests in CVS validate distributed GPU communication performance across AMD GPU clusters. The suites run RCCL collectives, optionally validate against expected thresholds, and generate HTML artifacts (graph and heatmap reports).

RCCL test suites#

CVS provides the following RCCL suites:

Test suite

What it does

rccl_perf

User-facing performance suite that runs configured collectives with environment script staging.

rccl_regression

Regression suite with Cartesian product sweep using environment variable combinations from JSON configuration.

All suites also collect host/network information and check firewall state before performance runs.

Run RCCL test commands#

See Run CVS RCCL performance and regression tests for cvs run examples, environment script staging, and heatmap generation.

Note

Update expected thresholds in results for your platform and cluster size before relying on pass/fail status. Placeholders such as {user-id} in config paths are resolved by CVS at runtime.

rccl_config.json#

Main config file used by rccl_perf and rccl_regression:

rccl_config.json (example)
{
  "rccl": {
    "env_source_script": "/home/{user-id}/thor2_env_script.sh",
    "mpi_params": {
      "no_of_nodes": "2",
      "no_of_local_ranks": "8",
      "mpi_pml": "auto",
      "mpi_dir": "/home/{user-id}/openmpi/bin",
      "mpi_oob_port": "eth0"
    },
    "rccl_test_params": {
      "rccl_collective": ["all_reduce_perf", "all_gather_perf", "broadcast_perf"],
      "rccl_tests_dir": "/home/{user-id}/rccl-tests/build",
      "start_msg_size": "1024",
      "end_msg_size": "16g",
      "step_function": "2",
      "warmup_iterations": "10",
      "no_of_iterations": "20",
      "data_types": ["float"]
    },
    "cvs_params": {
      "verify_bus_bw": "False",
      "verify_bw_dip": "True",
      "verify_lat_dip": "True",
      "cluster_snapshot_debug": "False",
      "rccl_result_file": "/home/{user-id}/rccl_result_file.json"
    },
    "results": {}
  }
}
rccl_config.json with regression (example)
{
  "rccl": {
    "env_source_script": "/home/{user-id}/thor2_env_script.sh",
    "mpi_params": {
      "no_of_nodes": "2",
      "no_of_local_ranks": "8",
      "mpi_pml": "auto",
      "mpi_dir": "/home/{user-id}/openmpi/bin"
    },
    "rccl_test_params": {
      "rccl_collective": ["all_reduce_perf"],
      "rccl_tests_dir": "/home/{user-id}/rccl-tests/build",
      "start_msg_size": "1024",
      "end_msg_size": "16g",
      "step_function": "2",
      "warmup_iterations": "10",
      "no_of_iterations": "20"
    },
    "regression": {
      "NCCL_ALGO": ["Ring", "Tree"],
      "NCCL_PROTO": ["Simple"],
      "NCCL_IB_QPS_PER_CONNECTION": ["1", "2"],
      "NCCL_PXN_DISABLE": ["0", "1"],
      "NCCL_MIN_NCHANNELS": ["8", "16"],
      "NCCL_MAX_NCHANNELS": ["8", "16"]
    },
    "cvs_params": {
      "verify_bus_bw": "False",
      "verify_bw_dip": "True",
      "verify_lat_dip": "True",
      "rccl_result_file": "/home/{user-id}/rccl_result_file.json"
    },
    "results": {}
  }
}

Parameters#

Configuration parameters for RCCL suites:

Core Configuration:

Configuration parameter

Example/default

Description

env_source_script

"/home/{user-id}/thor2_env_script.sh"

Environment script sourced before test execution (contains RCCL/NCCL/UCX tuning parameters).

rccl_collective

["all_reduce_perf", "all_gather_perf"]

RCCL collectives to execute.

rccl_result_file

"/tmp/rccl_result.json"

Output file path for RCCL parsed results.

start_msg_size

"1024"

Start message size for sweep.

end_msg_size

"16g"

End message size for sweep.

regression

{"NCCL_ALGO": ["Ring", "Tree"], "NCCL_PROTO": ["Simple"]}

Cartesian product sweep using NCCL/RCCL environment variable combinations (rccl_regression only).

MPI Parameters (mpi_params):

Configuration parameter

Example/default

Description

mpi_pml

"auto"

MPI point-to-point messaging layer (auto, ucx, ob1).

RCCL Test Parameters (rccl_test_params):

Configuration parameter

Example/default

Description

warmup_iterations

"10"

Warmup iteration count before measured iterations.

no_of_iterations

"20"

Number of measured iterations.

step_function

"2"

Message-size progression rule.

CVS Parameters (cvs_params):

Configuration parameter

Example/default

Description

verify_bus_bw

"False"

Enable bus-bandwidth threshold validation.

verify_bw_dip

"True"

Enable bandwidth-dip validation.

verify_lat_dip

"True"

Enable latency-dip validation.

cluster_snapshot_debug

"False"

Enables before/after cluster metric snapshots around tests.

results

{}

Expected threshold values used for pass/fail validation.

Expected results format#

The results section is used for threshold validation. Values are keyed by collective and message size (bytes), with expected bus bandwidth values.

results snippet
"results": {
  "all_reduce_perf": {
    "bus_bw": {
      "8589934592": "330.00",
      "17179869184": "350.00"
    }
  }
}

Collective meanings#

The following collectives can be specified in the results block; each name maps to a distinct RCCL operation.

  • all_reduce_perf: all ranks reduce then receive the reduced result.

  • all_gather_perf: each rank receives data from all ranks.

  • scatter_perf: root rank distributes shards to all ranks.

  • gather_perf: all ranks send data to a root rank.

  • reduce_scatter_perf: reduction followed by scatter.

  • sendrecv_perf: point-to-point pair communication.

  • alltoall_perf: equal-sized all-to-all exchange.

  • alltoallv_perf: variable-sized all-to-all exchange.

  • broadcast_perf: one-to-all data broadcast.

Regression testing#

The rccl_regression suite performs Cartesian product sweeps using the regression object in the configuration. Each key-value pair represents an NCCL/RCCL environment variable and its possible values:

"regression": {
  "NCCL_ALGO": ["Ring", "Tree"],
  "NCCL_PROTO": ["Simple"],
  "NCCL_IB_QPS_PER_CONNECTION": ["1", "2"],
  "NCCL_PXN_DISABLE": ["0", "1"],
  "NCCL_MIN_NCHANNELS": ["8", "16"],
  "NCCL_MAX_NCHANNELS": ["8", "16"]
}

Key features:

  • Cartesian product: All combinations of the environment variables are tested

  • Paired channels: NCCL_MIN_NCHANNELS and NCCL_MAX_NCHANNELS are paired (not Cartesian)

  • Tree filtering: Tree algorithm only runs with all_reduce_perf collective

  • Environment variables: Use actual NCCL/RCCL environment variable names as keys

Single-node testing#

Single-node RCCL testing can be achieved by configuring a single-node cluster in your cluster JSON file and running either rccl_perf or rccl_regression suites. The tests will automatically adapt to single-node execution when only one node is present in the cluster configuration.

Heatmap generation#

Generate performance heatmaps by comparing actual test results against a golden reference:

The following command generates an HTML heatmap that highlights per-collective bandwidth deviations from the reference.

cvs generate heatmap \
    --actual /tmp/rccl_perf_results.json \
    --reference /path/to/golden_reference.json \
    --output /var/www/html/cvs/rccl_heatmap.html \
    --title "RCCL Performance Comparison" \
    --metadata

Options:

  • --actual: Path to actual test results JSON file (required)

  • --reference: Path to golden reference JSON file (required)

  • --output: Output HTML file path (optional, defaults to /tmp)

  • --title: Custom heatmap title (optional)

  • --metadata: Include metadata table if actual JSON has ‘metadata’ key

  • --no-data-table: Exclude data table from output

Environment script setup#

All cluster configurations require an environment script to be sourced before RCCL tests, regardless of NIC type (Broadcom/ConnectX/AINIC):

Follow these steps to select and configure the correct environment script for your cluster hardware.

  1. For AINIC clusters: Ensure AMD ANP is installed and available on all target nodes, then edit input/config_file/rccl/ainic_env_script.sh and set ANP_HOME_DIR to your ANP install path.

  2. For other NIC types: Use the appropriate environment script (thor2_env_script.sh for Broadcom, cx7_env_script.sh for ConnectX-7, etc.).

  3. Set env_source_script in RCCL JSON config to the correct script path for your hardware.

  4. Run RCCL using cvs run ... commands; CVS sources the script automatically.

The environment script contains essential RCCL/NCCL/UCX tuning parameters and paths required for proper test execution.

Known issue: bnxt_re (Thor2) multi-node GPU Direct RDMA failure#

Symptom: On bnxt_re (Broadcom Thor2) clusters, multi-node RCCL/NCCL jobs fail symmetrically on every node during ncclCommInitRank with:

NET/IB: Peermem not available, falling back to GPU Direct RDMA (DMAbuf) for device 0
Call to ibv_reg_mr_iova2 failed with error Bad address

with a matching kernel-side error in dmesg:

infiniband bng_re0: bng_re_reg_user_mr: ib_umem_get failed! rc = -14

Root cause: This is not a per-node config or cabling issue – it reproduces identically on every node because it is a build/config gap. RCCL’s src/transport/net.cc only compiles its ROCm-native GPU memory registration path (hsa_amd_portable_export_dmabuf) when built against HIP_VERSION < 71260540. On newer ROCm builds (HIP_VERSION >= 71260540), it instead takes a CUDA-style code path that additionally requires ncclCuMemEnable() (NCCL_CUMEM_ENABLE) to be true. Because NCCL_CUMEM_ENABLE defaults to 0 on ROCm/RCCL, that condition is never met, so RCCL silently falls through to a plain (non-DMA-BUF) ibv_reg_mr_iova2 call on a raw GPU device pointer. Without a peer-memory kernel module, bnxt_re’s driver cannot pin that pointer via get_user_pages and rejects it with EFAULT (“Bad address”).

Fix: Set NCCL_CUMEM_ENABLE=1 (alongside NCCL_DMABUF_ENABLE=1) in the environment script used for bnxt_re/Thor2 clusters. input/env_file/rccl/thor2_env_script.sh sets this by default. Enabling CUMEM also lets NCCL recognize GPUs on a UALoE/scale-up fabric as directly P2P-reachable (P2P/CUMEMMNNVL), bypassing the NET/IB transport – and its DMA-BUF gap – entirely when a scale-up fabric is present.

Verified on a 2-node bnxt_re/Helios cluster (4 GPUs/node, 8 ranks total): with NCCL_CUMEM_ENABLE=0 every collective failed at ncclCommInitRank; with NCCL_CUMEM_ENABLE=1 all collectives (all_reduce_perf, all_gather_perf, reduce_scatter_perf, alltoall_perf, broadcast_perf; fp32 and bf16; 1 KB-1 GB) passed.

Validation and artifacts#

After a run, CVS produces the following outputs that can be used for pass/fail assessment and further analysis.

  • Test-level pass/fail is based on command execution plus enabled validations (results, bandwidth checks, dip checks).

  • Performance reports are generated under /tmp/rccl_perf_report_*.html.

  • Regression reports are generated under /tmp/rccl_perf_report_*.html with “RCCL Multi Node Performance Report” title.

  • Heatmaps are generated using cvs generate heatmap with customizable output paths.

  • All artifacts generated under /tmp can be copied to a web server path (for example /var/www/html/cvs) for browser access.