Run Cluster Validation Suite (CVS) test suites with a per-host Docker container backend#
2026-09-24
8 min read time
CVS can route workload commands through a long-lived per-host container instead of running them directly on the host filesystem. Use the container backend when you want to validate the same image you ship to production, keep the host footprint minimal (Docker, GPU driver, and SSH only), or pin the test environment byte-for-byte.
Note
Current scope: Of the test suites shipped with CVS, only rvs_cvs and install_rvs consume the orchestrator and honor the orchestrator key in the cluster file. All other cvs run test suites and the cvs exec CLI run on the host regardless of the orchestrator value. Migrating additional suites to the orchestrator is tracked separately. Custom Python scripts can use the OrchestratorFactory API directly as an escape hatch.
Prerequisites#
On every cluster node:
Docker installed. The SSH user needs either passwordless
sudo dockeror direct Docker access (for example membership in thedockergroup) – CVS probes once per run (sudo -n true) and caches which applies, prefixing every subsequent Docker command accordingly. See Cluster Validation Suite (CVS) cluster file: configuration and backend selection.Host driver loaded so
/dev/kfd,/dev/dri/*, and/dev/infiniband/*(when RDMA is in scope) are present for passthrough.SSH user home directory reachable. The orchestrator mounts
~/.sshas/host_sshinside the container so that the in-containersshdon port2224can authenticate.Container image either pre-loaded on every node (
docker load) or pullable from a reachable registry. The image must containopenssh-serverand the workload binaries the suite invokes (for example/opt/rocm/bin/rvs).
On the head node where you launch cvs run:
CVS installed (see Install Cluster Validation Suite (CVS) on ROCm).
SSH key-based access to every cluster node as the SSH user.
Step 1: Copy the cluster template#
CVS ships a cluster_container.json template alongside the baremetal cluster.json. Copy it to a working location:
cvs config copy cluster_container.json --output ~/cvs_workspace/cluster_container.json
You can browse every available template directory with:
cvs config list-dirs
Step 2: Edit the placeholders#
Replace every <changeme> placeholder in the copied file — CVS exits with an error if any placeholder remains unresolved. Key fields to set:
{user-id}: your SSH user (or leave it for runtime resolution).priv_key_file: absolute path to your SSH private key.head_node_dict.mgmt_ipand the keys ofnode_dict: real IPs or hostnames.mgmt_ipcan be one of thenode_dictkeys (usually the first node) or a completely separate host.container.image: an image present on every node or pullable from a reachable registry. The image must includeopenssh-serverand the workload binary (for examplervs).container.name: container name on each host. For parallel runs make this per-iteration unique (for examplecvs_iter_<run_id>). Pin it explicitly when usinglifetime: persistent.container.lifetime:no_launch,per_run, orpersistent. See the lifecycle note below.
For the full schema and runtime argument reference, see Cluster Validation Suite (CVS) cluster file: configuration and backend selection.
Step 3: Run a test suite#
Run rvs_cvs the same way you run any CVS test suite, but with the container cluster file:
cvs run rvs_cvs \
--cluster_file ~/cvs_workspace/cluster_container.json \
--config_file ~/cvs_workspace/health/mi300_health_config.json \
--html=/var/www/html/cvs/rvs.html \
--self-contained-html \
--capture=tee-sys \
--log-file=/tmp/rvs.log \
-vvv -s
What happens during the run:
The
orchfixture inrvs_cvsreadsorchestrator: containerfrom the cluster file and constructs aContainerOrchestrator.With
lifetime: per_run, CVS removes any container with the same name on each host, runs the configured image with the mergedruntime.args, and starts an in-containersshdon port2224.With
lifetime: persistent, CVS attaches to a container already running on every host, or starts one fresh if none is running.With
lifetime: no_launch, CVS verifies that a container with the configured name is already running on every host and reuses it.All
rvsinvocations are routed through the container viadocker execand the in-containersshd.
Step 4: Verify container execution#
Confirm that the workload ran inside the container, not on the host. From any node in the cluster, list running containers with the configured name:
ssh <user>@<node> sudo docker ps --filter name=^cvs_container$
You should see one running container per node with the configured name and image. You can fan this check out across every node from the head node:
cvs exec --cluster_file ~/cvs_workspace/cluster_container.json \
--cmd "sudo docker ps --filter name=^cvs_container$ --format '{{.Names}} {{.Image}}'"
Note
cvs exec always runs commands on the host regardless of the orchestrator value. That’s exactly what you want for verification: it lets you observe the host’s view of which containers are running.
Lifecycle and teardown#
The container.lifetime policy controls who owns the container lifecycle:
per_run(default) - CVS starts a fresh container at setup and force-removes it at teardown. Anything written to the container overlay is lost when the run ends.persistent- CVS attaches to the container if it is already running on every host, or starts it fresh if it is running on no host. Teardown is a no-op, so the container and its overlay survive across runs. Notes:Enables install-then-run workflows: run
cvs run install_rvsfollowed bycvs run rvs_cvsin separate invocations.If the container is running on some hosts but not all, CVS fails rather than rebuilding — rebuilding would destroy the overlay on the still-running hosts. Remove it on all hosts and rerun, or restart it on the missing hosts.
Pin
container.nameso a tag bump does not silently abandon the overlay.
no_launch- CVS does not start anything. It verifies that a container with the configured name is already running on every host and reuses it. Teardown is a no-op.
See the lifetime truth table in Cluster Validation Suite (CVS) cluster file: configuration and backend selection for the full state matrix.
To stop and remove a container that CVS left running (persistent or no_launch), run on every node:
cvs exec --cluster_file ~/cvs_workspace/cluster_container.json \
--cmd "sudo docker rm -f cvs_container"
Replace cvs_container with the actual container.name from your cluster file.
Common pitfalls#
Keep the following in mind when operating this feature.
Image without ``openssh-server``.
setup_sshdcannot startsshdon port2224andorch.execfails to connect. Make sure the image installsopenssh-serverand exposes the binary at/usr/sbin/sshd.Image without the workload binary.
cvs run rvs_cvsinvokesrvsinside the container; if the image lacks/opt/rocm/bin/rvsthe test fails with a “command not found” style error.``lifetime: persistent`` without a pinned ``name``. The default container name is
<user>_<sanitized_image>, which shifts when you bump the image tag. A tag bump silently abandons the previous container’s overlay (installs, clones) and starts fresh. Pincontainer.namewhen usingpersistent.Port ``2224`` already bound on the host. With
network: hostthe in-containersshdbinds to2224in the host’s network namespace. If something else on the host already listens on2224the bind fails. Stop the conflicting service.Picked the wrong cluster template. Use
cluster_container.json(withorchestrator: container) for container mode. Usingcluster.jsonruns the suite on the host even when the image you wanted to test is sitting on every node.