Run ad-hoc commands across all Cluster Validation Suite (CVS) cluster nodes via SSH#
2026-09-24
8 min read time
CVS provides an exec command to execute arbitrary shell commands on all nodes in the cluster simultaneously using parallel SSH. This is useful for gathering system information, running diagnostics, or performing administrative tasks across all cluster nodes at once.
The exec command uses the same cluster configuration files as other CVS commands and supports both command-line arguments and environment variables for configuration.
Execute commands#
To execute a command on all cluster nodes:
cvs exec --cmd "hostname" --cluster_file ~/cvs_workspace/cluster.json
You can also use the CLUSTER_FILE environment variable:
CLUSTER_FILE=~/cvs_workspace/cluster.json cvs exec --cmd "hostname"
Command options#
The exec command supports these options:
--cmd: The shell command to execute on all nodes (required)--cluster_file: Path to cluster configuration JSON file (optional ifCLUSTER_FILEenvironment variable is set)--target: Scope of execution —computes(default),switches, orall--timeout: Per-node command output timeout in seconds (default:30). Controls how long to wait for each host’s stdout after the SSH connection is established.--connect-timeout: Per-node SSH connection timeout in seconds (default:15). Unreachable hosts fail fast after this many seconds regardless of--timeout.--json: Emit results as a single JSON object instead of human-readable text. Useful for scripting and piping tojq.--verbose/-v: Show internal SSH diagnostics (SocketDisconnectError,AuthenticationError, pruning messages). Suppressed by default to keep output clean.
Target scope#
By default cvs exec runs on every host in node_dict (compute nodes). Use --target to change the scope:
|
Hosts targeted |
|---|---|
|
All entries in |
|
All |
|
Both compute nodes and switch trays |
# Compute nodes only (default — identical to old behavior)
cvs exec --cmd "hostname" --cluster_file ~/cvs_workspace/cluster.json
# Switch trays only
cvs exec --cmd "show version" --cluster_file ~/cvs_workspace/cluster_rack.json --target switches
# Both compute nodes and switch trays
cvs exec --cmd "date" --cluster_file ~/cvs_workspace/cluster_rack.json --target all
Rack-aware cluster file#
--target switches and --target all require a racks block in the cluster file that lists the switch tray IPs and (optionally) per-rack SSH credentials. Copy the rack template with cvs config copy cluster_rack.json --output ~/cvs_workspace/cluster_rack.json.
If --target switches is used with a plain cluster file (no racks block), CVS prints a warning and exits without executing any command.
Exec command examples#
The following examples show common diagnostic and information-gathering commands.
System information#
# Check system uptime on all nodes
cvs exec --cmd "uptime" --cluster_file ~/cvs_workspace/cluster.json
# Get OS and kernel information
cvs exec --cmd "uname -a" --cluster_file ~/cvs_workspace/cluster.json
# Check available memory
cvs exec --cmd "free -h" --cluster_file ~/cvs_workspace/cluster.json
GPU information#
# Get GPU information using rocm-smi
cvs exec --cmd "rocm-smi --showid --showproductname --showuniqueid" --cluster_file ~/cvs_workspace/cluster.json
# Get GPU temperature and usage
cvs exec --cmd "rocm-smi --showtemp --showuse" --cluster_file ~/cvs_workspace/cluster.json
# Get GPU memory information
cvs exec --cmd "rocm-smi --showmeminfo vram" --cluster_file ~/cvs_workspace/cluster.json
Network information#
# List RDMA devices
cvs exec --cmd "rdma res" --cluster_file ~/cvs_workspace/cluster.json
# Check network interfaces
cvs exec --cmd "ip addr show" --cluster_file ~/cvs_workspace/cluster.json
# Check network connectivity
cvs exec --cmd "ping -c 3 8.8.8.8" --cluster_file ~/cvs_workspace/cluster.json
# Show routing table
cvs exec --cmd "ip route show" --cluster_file ~/cvs_workspace/cluster.json
Storage and disk information#
# Check disk usage
cvs exec --cmd "df -h" --cluster_file ~/cvs_workspace/cluster.json
# Check mounted filesystems
cvs exec --cmd "mount | grep -E '(ext4|xfs|nfs)'" --cluster_file ~/cvs_workspace/cluster.json
# Check disk I/O statistics
cvs exec --cmd "iostat -x 1 3" --cluster_file ~/cvs_workspace/cluster.json
Process and service information#
# Check running processes
cvs exec --cmd "ps aux | head -20" --cluster_file ~/cvs_workspace/cluster.json
# Check system load
cvs exec --cmd "top -b -n 1 | head -10" --cluster_file ~/cvs_workspace/cluster.json
# Check systemd services status
cvs exec --cmd "systemctl status rocm" --cluster_file ~/cvs_workspace/cluster.json
Output format#
Each result is prefixed with the target type (compute or switch) and the host IP or hostname, making it easy to identify which output came from which node:
[compute] Host: 10.0.0.2
Linux node01 5.15.0-91-generic #101-Ubuntu SMP Tue Nov 14 13:30:08 UTC 2023 x86_64 x86_64 x86_64 GNU/Linux
---
[compute] Host: 10.0.0.3
Linux node02 5.15.0-91-generic #101-Ubuntu SMP Tue Nov 14 13:30:08 UTC 2023 x86_64 x86_64 x86_64 GNU/Linux
---
When --target all is used, switch tray output follows the compute output with a [switch] prefix:
[switch] Host: 192.168.1.1
SONiC Software Version: SONiC.HEAD.xxx
---
This format helps you quickly distinguish compute and switch output when troubleshooting or gathering information across the cluster.
JSON output#
Pass --json to receive a single structured JSON object on stdout. All SSH diagnostics are automatically suppressed so the output is pipe-safe.
cvs exec --cmd "hostname" --cluster_file ~/cvs_workspace/cluster.json --json
Output schema:
{
"command": "hostname",
"read_timeout": 30,
"connect_timeout": 15,
"output": {
"10.0.0.2": "node01\n",
"10.0.0.3": "node02\n",
"10.0.0.4": "ABORT: Host Unreachable Error"
}
}
outputis a flathost → stringmap. Error and timeout strings appear as values for unreachable hosts.For
--target all, compute and switch hosts are merged into the sameoutputdict.
Pipe directly to jq for filtering:
# Print only the output dict
cvs exec --cmd "hostname" --json | jq '.output'
# Get a single host's output (use bracket notation — IPs contain dots)
cvs exec --cmd "hostname" --json | jq -r '.output["10.0.0.2"]'
# List only hosts that succeeded (no error strings)
cvs exec --cmd "hostname" --json | jq '[.output | to_entries[] | select(.value | test("ABORT|Error") | not) | .key]'
Timeout behavior#
--timeout and --connect-timeout control two distinct phases of the per-node SSH operation:
Flag |
Phase controlled |
Default |
|---|---|---|
|
TCP/SSH handshake per host |
|
|
Command output reading after connection |
|
For short diagnostic commands, set both to the same value to ensure unreachable hosts fail within that window:
cvs exec --cmd "uptime" --timeout 10 --connect-timeout 10
For long-running operations (benchmarks, firmware updates), keep --connect-timeout small and raise --timeout only:
cvs exec --cmd "./run_benchmark.sh" --timeout 600 --connect-timeout 10
Note
When TCP packets to target hosts are silently dropped (for example, no VPN/sshuttle tunnel), the effective per-host timeout is governed by --timeout (the socket-level IO timeout), not --connect-timeout. Set --timeout equal to --connect-timeout for the fastest failure in that scenario.
Troubleshooting#
If you encounter connection issues:
Verify your SSH credentials in the cluster configuration file
Ensure SSH key-based authentication is properly set up
Check network connectivity to all cluster nodes
Verify the cluster file format and paths
If commands fail on specific nodes:
Check if the command exists on those nodes
Verify user permissions for executing the command
Check if required packages or tools are installed on all nodes
If --target switches produces a warning and no output:
Verify that the cluster file contains a
racksblock withswitch_traysentriesCopy a rack-aware cluster file with
cvs config copy cluster_rack.json --output ~/cvs_workspace/cluster_rack.json
If SSH connections hang or time out:
Reduce
--connect-timeout(default15s) to fail unreachable hosts fasterReduce
--timeout(default30s) if commands should complete quicklyConfirm network reachability to all target hosts before running (for example, via sshuttle or VPN)
If the output contains unexpected SSH diagnostic messages:
These are suppressed by default; they appear only when
--verbose/-vis passedIf they still appear, check that the root logger level has not been overridden elsewhere
Tip
Use --json | jq '.output' to get a clean, machine-readable summary. Combine with --verbose only when debugging connection issues.
Next steps#
Copy files to all cluster nodes using Cluster Validation Suite (CVS) parallel SCP — copy files or directories to all cluster nodes in parallel using
cvs scp.Run Cluster Validation Suite (CVS) test suites on AMD Instinct GPU clusters — run a test suite against the cluster.
Cluster Validation Suite (CVS) cluster file: configuration and backend selection — full cluster file schema including the
racksblock required for--target switches.