Monitor AMD Instinct GPU cluster health and metrics with Cluster Validation Suite (CVS)

Monitor AMD Instinct GPU cluster health and metrics with Cluster Validation Suite (CVS)#

2026-09-24

2 min read time

Applies to Linux

CVS provides three monitoring paths:

  • Health reports — run cvs monitor check_cluster_health for a point-in-time HTML health report over SSH (no agents on nodes).

  • Live dashboards — real-time and historical metrics while workloads run:

    • Agentless (SSH) — CVS Cluster Monitor dashboard (cvs/monitors/cluster-mon/)

    • Agent (Exporters) — Prometheus/Grafana fleet monitor (cvs/monitors/metrics_exp/)

Choose health reports for preflight validation or triage snapshots. Choose live dashboards for ongoing visibility during training or inference campaigns.