Live AMD Instinct GPU cluster dashboards with Cluster Validation Suite (CVS)#
2026-09-24
2 min read time
Live dashboards keep GPU and network metrics visible while workloads run—real-time views and historical trends in Grafana or the CVS Cluster Monitor UI. CVS ships two approaches under cvs/monitors/:
The following options are available.
Approach |
How it works |
Best for |
|---|---|---|
Dashboard polls nodes over SSH ( |
Quick dashboard, minimal node footprint, MI300/MI325 clusters with SSH access |
|
Installs Prometheus exporters on nodes; Fleet Monitor UI + Grafana + Loki |
Large fleets, historical retention, Slurm/K8s control-plane views |
Both support jump hosts and parallel SSH. Choose agentless when you cannot install services on compute nodes; choose exporters when you need Prometheus retention, alerting, and multi-week trends.
For a one-time HTML health report (no dashboard), use Generate AMD Instinct GPU cluster health reports with Cluster Validation Suite (CVS).