Generate AMD Instinct GPU cluster health reports with Cluster Validation Suite (CVS)

Generate AMD Instinct GPU cluster health reports with Cluster Validation Suite (CVS)#

2026-09-24

3 min read time

Applies to Linux

Health reports run cvs monitor check_cluster_health from the head node. CVS SSHs to each compute node, collects GPU/NIC counters and logs, and writes a self-contained HTML report. No agents or exporters are installed on the cluster.

The monitor identifies hardware degradation (RAS, PCIe/XGMI, RDMA counters via AMD SMI) and software failures (dmesg and journalctl signatures). Use it before a test campaign, after a change, or to snapshot counters while workloads run and compare deltas across nodes.

Generate a health report#

  1. Complete Install CVS and Set up a cluster file.

  2. List available monitors:

    cvs monitor
    
  3. View help:

    cvs monitor check_cluster_health --help
    
  4. Run the monitor:

    Option

    Description

    --cluster_file PATH

    Cluster JSON (same file as cvs run / cvs exec). Takes precedence over CLUSTER_FILE.

    --iterations N

    Number of sampling iterations

    --time_between_iters SECONDS

    Sleep between iterations

    --report_file PATH

    Output HTML path (default: cluster_report.html in the current directory)

    Example:

    cvs monitor check_cluster_health \
      --cluster_file ~/cvs_workspace/cluster.json \
      --iterations 2
    

    Or set CLUSTER_FILE once:

    export CLUSTER_FILE=~/cvs_workspace/cluster.json
    cvs monitor check_cluster_health --iterations 2
    

    Note

    Deprecated flags --hosts_file, --username, --key_file, and --password still work; prefer --cluster_file for consistency with other CVS commands.

  5. Open cluster_report.html in a browser.

Review the health report#

The report includes GPU and NIC snapshots, historic error logs, and per-iteration deltas for triage. Anomalies are highlighted in tables (PCIe, RDMA, GPU errors, cable issues, kernel logs).

../../../_images/rdma.png ../../../_images/pcie.png ../../../_images/journlctl.png

Metrics are collected with AMD SMI, ethtool, and rdma utilities on each node. See the AMD SMI CLI reference for field definitions.

For live dashboards instead of one-off HTML reports, see Live AMD Instinct GPU cluster dashboards with Cluster Validation Suite (CVS).