Agentless SSH live dashboards with the Cluster Validation Suite (CVS) Cluster Monitor#
2026-09-24
4 min read time
The CVS Cluster Monitor (cvs/monitors/cluster-mon/) is a live dashboard that polls the cluster over SSH. It does not install exporters on GPU nodes — collectors run amd-smi, RDMA, and log queries remotely on a configurable interval.
Use this when you want real-time GPU/NIC views, heatmaps, topology, and logs without deploying Prometheus exporters on every host.
Features#
This approach provides the following features.
Real-time GPU metrics: utilization, temperature, power, memory, PCIe, ECC, XGMI
Network: RDMA statistics, LLDP topology, NIC firmware and driver info
Logs: filtered
dmesg, journal errors, userspace failures, custom grep searchTCP reachability probes before SSH (fast skip of unreachable nodes)
Jump host support and web-based configuration (nodes, SSH keys, polling interval)
Prerequisites#
The following prerequisites are required.
Docker and Docker Compose v2 on the monitoring host
SSH access to cluster nodes (direct or via jump host)
amd-smion GPU nodes; RDMA tools optional for network views
Docker Compose also starts a Redis sidecar for metric snapshots and events (no separate Redis install on the host).
Quick start (Docker)#
From your CVS checkout, go to
cvs/monitors/cluster-mon/.Prepare configuration under
config/:cp config/cluster.yaml.example config/cluster.yaml cp config/nodes.txt.example config/nodes.txt
Edit
cluster.yaml(SSH user, key path, polling interval, optional jump host) andnodes.txt(one host per line). Seecvs/monitors/cluster-mon/README.mdfor the full schema.Build and deploy (recommended):
./full-rebuild.sh
The script builds the image, runs
docker-compose up -d, seeds config from the examples if missing, and triggers a config reload so monitoring starts automatically.Alternatively, after editing
config/manually:docker-compose up -d --build
Open the dashboard at
http://<monitor-host>:8005(host port 8005 maps to the app inside the container on port 8001).Optional — change settings without rebuilding: open the Configuration tab, upload SSH keys if needed, edit nodes or jump-host settings, then Save Configuration and Start Monitoring.
Verify deployment#
docker-compose logs -f
curl http://<monitor-host>:8005/health
The health endpoint reports collection status (for example ssh_manager, collecting, connected clients).
Operational notes#
Keep the following in mind when operating this feature.
Default metrics interval: 60 seconds (
polling.intervalincluster.yamlorPOLLING__INTERVALenv var). For large fleets (50+ nodes), consider 120 seconds.Host reachability is re-probed every 5 minutes; SSH clients refresh when nodes come back online
Stop or restart:
docker-compose down/docker-compose restartLLDP packages can be installed cluster-wide from the Configuration tab
Other deployment options#
cvs/monitors/cluster-mon/DEPLOYMENT.md covers bare-metal (Python backend + React frontend), Nginx reverse proxy, systemd, resource limits, upgrades, and troubleshooting. Note: prefer docker-compose and port 8005 on the host — some older examples in that file reference port 8001 on the host.
For Prometheus/Grafana-based monitoring with exporters installed on nodes, see Live GPU fleet dashboards with Cluster Validation Suite (CVS) Prometheus exporters.