Live AMD Instinct GPU cluster dashboards with Cluster Validation Suite (CVS)

Live AMD Instinct GPU cluster dashboards with Cluster Validation Suite (CVS)#

2026-09-24

2 min read time

Applies to Linux

Live dashboards keep GPU and network metrics visible while workloads run—real-time views and historical trends in Grafana or the CVS Cluster Monitor UI. CVS ships two approaches under cvs/monitors/:

The following options are available.

Approach

How it works

Best for

Agentless (SSH)

Dashboard polls nodes over SSH (amd-smi, RDMA, logs); no exporters on workers

Quick dashboard, minimal node footprint, MI300/MI325 clusters with SSH access

Agent (Exporters)

Installs Prometheus exporters on nodes; Fleet Monitor UI + Grafana + Loki

Large fleets, historical retention, Slurm/K8s control-plane views

Both support jump hosts and parallel SSH. Choose agentless when you cannot install services on compute nodes; choose exporters when you need Prometheus retention, alerting, and multi-week trends.

For a one-time HTML health report (no dashboard), use Generate AMD Instinct GPU cluster health reports with Cluster Validation Suite (CVS).