Cluster Validation Suite (CVS) quickstart: install and run your first cluster command

Cluster Validation Suite (CVS) quickstart: install and run your first cluster command#

2026-09-24

5 min read time

Applies to Linux

This guide gets you from zero to your first cluster-wide cvs exec in about 15 minutes.

Prerequisites#

The following prerequisites are required before you begin:

Note

Head node: The Linux host where you install and run the CVS CLI. It can be a VM or bare metal and does not need a GPU. It must be able to SSH to every worker.

The head node can be either of the following:

  • One of the cluster members in node_dict — usually the first node.

  • A completely separate host that is not in node_dict.

Set head_node_dict.mgmt_ip in cluster.json to that host’s address. See Cluster Validation Suite (CVS) cluster file: configuration and backend selection.

Step 1: Install CVS#

Clone the repository and install:

git clone https://github.com/ROCm/cvs
cd cvs
make install
source .cvs_venv/bin/activate
cvs --version
cvs list

Step 2: Copy the cluster file#

Every CVS command needs a cluster.json that lists your nodes and SSH credentials.

mkdir -p ~/cvs_workspace
cvs config copy cluster.json --output ~/cvs_workspace/cluster.json

Edit cluster.json and replace every <changeme> placeholder with your node hostnames, SSH user, and key path. CVS exits with an error if any placeholder remains unresolved. See Configure the Cluster Validation Suite (CVS) cluster file (cluster.json) and Cluster Validation Suite (CVS) cluster file: configuration and backend selection.

Step 3: Run cluster-wide commands#

Use cvs exec to run a shell command on every node in parallel:

cvs exec --cmd "hostname" --cluster_file ~/cvs_workspace/cluster.json

Successful output looks like this — one block per node:

[compute] Host: 10.0.0.2
node01
---
[compute] Host: 10.0.0.3
node02
---

You should see one hostname per node in your cluster. You can also set CLUSTER_FILE once and omit --cluster_file on later commands. See Run ad-hoc commands across all Cluster Validation Suite (CVS) cluster nodes via SSH for --target, --json, and timeouts.

Validate success#

If cvs exec returned a hostname from every node in your cluster file, CVS is installed and connected. Your cluster is ready to run tests.

If any node is missing from the output, check:

  • Passwordless SSH works from the head node to that node: ssh <user>@<host> hostname.

  • The node’s hostname or IP in cluster.json is correct and reachable.

  • The SSH key path and user in cluster.json match what works in the manual SSH check above.

Run the GPU visibility check to confirm AMD GPUs are visible on all nodes before running any test suite:

cvs exec --cmd "amd-smi list" \
  --cluster_file ~/cvs_workspace/cluster.json

Successful output from each node looks like this:

[compute] Host: 10.0.0.2
GPU: 0
    BDF: 0000:03:00.0
    UUID: c30074a5-0000-1000-81ab-4f2e7c6d90b1
    KFD_ID: 42109
    NODE_ID: 2
    PARTITION_ID: 0
<<truncated>>
---
[compute] Host: 10.0.0.3
GPU: 0
    BDF: 0000:23:00.0
    UUID: a10074a5-0000-1000-8042-6e1c8a9b52d0
    KFD_ID: 31758
    NODE_ID: 4
    PARTITION_ID: 0
<<truncated>>
---

Every node should report its GPU devices. A node that returns no GPUs or an error indicates a driver or device access issue to resolve before testing.

What to do next#

CVS is installed and your cluster is connected. The next step is to run a test suite.